$50–100/hr · Mercor · Hourly
A French voice actor who records high-quality speech samples in diverse styles to train AI text-to-speech systems for customer service applications.
What you would do
- Record voice samples across conversational, narrative, instructional, and other script types in standard, neutral French
- Deliver expressive speech with deliberate control over tone, pace, emphasis, and emotional nuance
- Execute multiple takes of the same script with variations in emotion, style, or inflection as directed
- Maintain vocal consistency, accent uniformity, and pronunciation precision throughout extended recording sessions
- Follow technical recording specifications for microphone placement, audio file formatting, and delivery quality
Who they want
- Native or native-level French speaker with standard, neutral accent (no strong regional dialect or foreign influence)
- Clear, expressive voice with strong control over intonation, diction, and vocal dynamics across emotional ranges
- Professional or near-professional recording setup including quality microphone, quiet environment, and pop filter
- Prior experience with voiceover, audiobook narration, TTS recording, or broadcast work is preferred but not required
- Availability for approximately 4-hour recording session with willingness to consent to voice cloning for AI training
Main skills
What the interview asks about
1.Neutral accent and language consistency
AI voice cloning depends on clean, unaccented pronunciation; regional variations or non-standard speech patterns degrade model training and the cloned voice's quality.
For example: “Describe your French accent and background. Born in Quebec, Paris, or Belgium? How would you deliver 'bonjour, comment puis-je vous aider?' to sound standard and neutral?”
2.Emotional delivery and vocal variation
Customer service AI needs to express different emotional contexts; without genuine vocal range and the ability to hit specific emotional tones on command, TTS output sounds robotic.
For example: “You record the line 'I understand your frustration.' Record versions conveying empathy, calm professionalism, and urgency - what changes in your voice, and how would you execute three distinct versions within one session?”
3.Script fidelity with natural delivery
Scripts must be read exactly as written for consistency, but obvious script-reading ruins authenticity; TTS models trained on unnatural delivery produce awkward synthetic speech.
For example: “A script says 'Thank you for your patience. We are currently experiencing high call volume.' How do you make this sound like you're talking to a real person rather than reading, while hitting every word and phrase exactly as written?”
4.Recording technical execution
Poor audio quality (background noise, clipping, inconsistent levels, equipment artifacts) renders recordings unusable for voice cloning or adds noise to training data.
For example: “Your recording setup: microphone type/placement, settings, noise control, monitoring. First 20 minutes have subtle hum. Do you re-record those takes and how do you prevent it for the final 3 hours?”
5.Vocal stamina and session consistency
A 4-hour recording session requires the ability to maintain voice quality, consistency, and emotional precision; fatigue causes the later recordings to deteriorate in quality or shift in accent.
For example: “You're 3 hours into a 4-hour session. Your voice feels fatigued, and you notice your pronunciation drifting slightly (vowels getting less crisp). What do you do to recover quality for the final hour of recordings?”
A task you may get
Record a 1-2 minute script in French across three distinct emotional contexts (professional-neutral, warm-empathetic, friendly-conversational). Demonstrate consistency in accent and pronunciation while showing intentional variation in tone and delivery.
How to prepare
- Research standard French pronunciation and identify your own regional influences; practice neutralizing any non-standard accent patterns
- Study examples of professional French voiceover work, TTS systems, and customer service voice acting to understand the expected delivery style
- Test your recording setup and environment in multiple locations and times to identify and eliminate background noise and audio artifacts
- Prepare a variety of script types (short sentences, longer narratives, dialogue) and practice recording them with different emotional intentions to build vocal flexibility
The facts
- Pay
- $50–100/hr
- Hours
- Hourly
- Where
- Remote
- Field
- Language and Audio
- Posted
- 6/19/2026
- Places left
- 13
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.