$50–100/hr · Mercor · Hourly
You provide expressive voice recordings of varied scripts to serve as training data for advanced AI voice synthesis.
What you would do
- Record voice samples reading conversational, narrative, and instructional scripts with natural pacing and authentic emotional range
- Perform multiple takes of the same content with varying emphasis, emotion, and vocal style as directed
- Maintain consistency in accent, tone, and voice quality throughout the 4-hour session to ensure AI trains on a coherent voice identity
- Follow technical recording setup guidelines including microphone placement, file naming, and audio format requirements
- Adjust delivery based on director feedback between takes to ensure recorded variations serve the AI training goal
Who they want
- USA-based voice actor with natural African American English delivery and clear articulation
- Natural, expressive voice with strong control over intonation, pacing, and emotional tone without formal training required
- Professional or semi-professional recording setup including quality microphone, acoustically treated space, and pop filter
- Comfort with voice cloning technology and explicit understanding that your voice recordings will be used to create synthetic versions for the client's AI agent
- Flexibility to deliver multiple vocal styles and emotion-inflected performances across consecutive takes within a 4-hour window
What the interview asks about
1.Maintaining voice consistency across styles
AI voice cloning requires consistent acoustic features; a voice that shifts dramatically between takes confuses the training process and produces artifacts.
For example: “You're recording the same sentence five times: once conversational, once corporate, once energetic, once calm, once neutral. How do you keep your core voice identity constant while changing emotional register?”
2.Responding to technical direction
Voice coaches and AI engineers have specific needs for the training data (e.g., 'less breathiness,' 'slower pacing on this phrase'); you must execute adjustments precisely.
For example: “The director says, 'On take three, you dropped the ending consonants. On take four, your intonation went up at the wrong place. Take five should combine the pacing of take three with the intonation curve of take two.' How do you execute this?”
3.Recording environment quality management
Background noise, room reflections, or mic technique flaws create training data noise that degrades the final AI voice.
For example: “Halfway through the session, you notice a subtle refrigerator hum that wasn't obvious at the start. How do you diagnose whether it's in your recordings, and what do you do about previous takes?”
4.Emotional nuance without overacting
AI needs authentic vocal emotion, not theatrical overdone delivery; recognizing the boundary is critical to creating usable training data.
For example: “A script is a customer service response to an angry customer. Show concern and patience, but don't sound condescending or melodramatic. How do you find that middle ground on take one versus take three?”
A task you may get
Record a 30-second script three different ways (e.g., conversational, formal, excited) maintaining consistent voice while shifting emotional tone, ensuring clear audio quality.
How to prepare
- Test your recording setup (microphone, room acoustics) by recording 5 minutes of natural speech and listening for background noise, plosives, or technical issues
- Practice scripts aloud multiple times focusing on different emotional registers without changing your natural accent or core voice identity
- Familiarize yourself with audio file formats and naming conventions (likely 16-bit WAV or MP3) to ensure submission-ready recordings
The facts
- Pay
- $50–100/hr
- Hours
- Hourly
- Where
- Remote
- Open to
- USA, USA
- Field
- Language and Audio
- Posted
- 7/17/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.