$50–100/hr · Mercor · Hourly
A Northern England-based female voice performer with authentic regional accent who records voice samples for text-to-speech AI model training.
What you would do
- Record high-quality voice samples from scripts spanning conversational, narrative, instructional, and various tonal contexts
- Deliver clear, expressive speech with precise control over tone, pacing, and pronunciation throughout takes
- Perform multiple versions of scripts with deliberate variations in emotion, emphasis, and vocal style
- Maintain consistent voice, accent, and delivery quality across multiple recording sessions
Who they want
- Female voice performer with natural, authentic Northern English accent (Yorkshire, Manchester, Newcastle, Liverpool region)
- Currently based in northern UK
- Clear, natural, expressive voice with strong emotional range and diction control
- Access to professional or near-professional recording setup with quality microphone, pop filter, and quiet environment
- Availability for approximately 4-hour recording session, possibly followed by additional session
What the interview asks about
1.Accent authenticity maintenance
AI models training on regional TTS need genuine accent characteristics without artificial exaggeration that would make output sound unconvincing.
For example: “A script included regional vocabulary and phrasing characteristic of Yorkshire English. How would you deliver that authentically while ensuring clarity for audiences unfamiliar with the region?”
2.Emotional range and consistency
TTS models improve when trained on the same voice across different emotional contexts, and models need consistency to learn voice characteristics.
For example: “You recorded the same sentence in five emotional contexts: neutral, enthusiastic, frustrated, calm, and urgent. How did you shift your delivery between takes while keeping your voice recognizable?”
3.Directional follow-through
Recording for AI requires precision in following technical direction about pacing, emphasis, and style to ensure usable training data.
For example: “A script required you to emphasize different words across three takes: emphasize 'really' in version A, 'problem' in version B, and 'solution' in version C. How did you adjust your delivery while maintaining natural speech flow?”
4.Technical recording quality
Poor audio quality compromises the entire training dataset, so understanding and maintaining technical standards is essential to the work.
For example: “You noticed background noise from traffic during recording. How would you adjust your setup or technique to maintain required audio quality without compromising take consistency?”
5.Script interpretation without over-rehearsal
Voice models need natural delivery that follows scripts precisely but doesn't sound over-produced or artificially polished.
For example: “A corporate script felt stilted when read straightforwardly but needed to sound warm and natural. How would you find a delivery that serves the script's intent without sounding wooden?”
A task you may get
Record three versions of the same brief conversational script: neutral delivery, enthusiastic delivery, and cautious delivery. Maintain consistent voice characteristics across all versions while clearly varying emotional tone.
How to prepare
- Practice recording yourself reading various scripts with different emotional framings, noting which techniques help vary delivery while maintaining voice consistency
- Review your recording setup: check microphone quality, background noise levels, and file format capabilities
- Record sample clips in your natural Northern accent and listen critically for clarity and regional authenticity
- Study examples of professional TTS voice recordings to understand clarity and naturalness standards
The facts
- Pay
- $50–100/hr
- Hours
- Hourly
- Where
- Remote
- Open to
- GBR, GBR
- Field
- Language and Audio
- Posted
- 6/1/2026
- Places left
- 15
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.