$50–150/hr · Mercor · Hourly
You record voice samples for AI text-to-speech systems, providing diverse emotional readings and vocal styles for training conversational AI agents.
What you would do
- Record scripted material across conversational, narrative, and instructional contexts for AI training
- Perform multiple takes with different emotional colorations and delivery styles
- Maintain consistent vocal qualities and American English accent throughout the session
- Follow precise technical specifications for audio recording environment and file formatting
Who they want
- Native American English speaker currently residing in the United States
- Access to professional or near-professional recording equipment and quiet recording space
- Strong command of vocal pacing, intonation, and emotional expression in speech
- Comfort with voice cloning technology applied to your recordings
Main skills
What the interview asks about
1.Emotional expression and vocal versatility
AI agents need training data showing how tone, pace, and emphasis convey different emotional states. Your ability to execute diverse emotional readings affects model quality.
For example: “You're recording a customer service greeting for both empathetic support and energetic sales. How would you technically approach recording the same script with distinct emotional flavors while maintaining the same accent?”
2.Technical audio quality management
Professional recording quality directly impacts whether AI training data is usable. Your setup and technique during four hours of recording determine whether the dataset meets technical standards.
For example: “Your session involves two hours of intimate conversational tone, then two hours of formal instructional content. What technical adjustments in microphone position would you make to capture both styles with consistent audio?”
3.Script precision with natural inflection
AI models learn from exact word delivery while needing realistic vocal patterns. Finding the balance between script fidelity and natural speech is critical for agent authenticity.
For example: “A script reads: 'Your account has been updated.' How would you deliver this exact wording multiple times with different emphasis patterns so AI learns when to prioritize different words?”
4.Consistency in long recording sessions
AI models trained on a single voice must hear consistent vocal characteristics across hundreds of utterances. Vocal drift significantly impacts model quality.
For example: “Four hours into an intense session with multiple vocal shifts, the director asks for five more takes of a script you recorded at hour one. How would you ensure your voice characteristics match your original recordings?”
5.Adapting to direction and multiple takes
Recording for AI requires understanding why variations matter. Your ability to execute specific adjustments based on feedback shows professional capability.
For example: “After 30 minutes, the director says: 'Slow your pacing by 20 percent and increase emphasis on technical terms.' How would you implement these adjustments across the remaining three and a half hours?”
6.Voice cloning consent and implications
You're authorizing your voice characteristics to be extracted and synthesized artificially. Understanding this affects your vocal delivery choices.
For example: “Knowing AI will clone your vocal patterns, how does this affect which vocal quirks or mannerisms you would use in your recordings?”
A task you may get
Record a 60-second customer service greeting with three distinct emotional variations - formal, warm/empathetic, and energetic - maintaining exact word accuracy while shifting vocal tone and pacing.
How to prepare
- Listen to professional voice actor work in customer service AI systems and note vocal techniques
- Test your recording setup with sample reads to identify technical audio issues
- Practice recording the same sentence with different emotions to find your range limits
- Research how voice cloning technology uses vocal characteristics to understand what consistency matters
The facts
- Pay
- $50–150/hr
- Hours
- Hourly
- Where
- Remote
- Field
- Language and Audio
- Posted
- 3/20/2026
- Places left
- 29
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.