$50/hr · Mercor · Hourly
You record high-quality voice samples in Canadian French for AI text-to-speech training, with control over tone, pacing, and emotional content.
What you would do
- Record voice samples from diverse script types: conversational, narrative, instructional, emotional
- Deliver fluid, natural speech with excellent command of tone, pacing, and pronunciation
- Record multiple renderings with variation in emotion, emphasis, and style per guidance
- Maintain consistent voice, accent, and delivery across recording sessions
- Follow detailed recording guidelines for environment, microphone setup, and file formatting
Who they want
- Canadian French speaker based in Canada with native fluency
- Ability to speak international French with neutral accent
- Well-articulated, expressive delivery (prior voice work, narration, or broadcasting is a plus)
- Professional or near-professional recording setup with quality microphone and quiet environment
- Availability for a single 4-hour session with possibility of a follow-up session
Main skills
What the interview asks about
1.Canadian French versus international French
The role requires both native Canadian French authenticity and the ability to shift to neutral international French; this tests that bilingual flexibility.
For example: “A script uses colloquialisms common to Quebec French. The AI product needs neutral French. Would you record the script as written, adapt it, or ask for clarification? What's at stake?”
2.Recording environment and technical setup
Audio quality directly affects AI training; this tests whether you understand what constitutes professional-grade recording conditions.
For example: “You've set up recording in your home office with a quality microphone and pop filter. When you review your first take, you hear background HVAC noise. What would you do?”
3.Expressive delivery with control
TTS systems learn variety from multiple takes with different emotional colorings; this tests whether you can consciously vary delivery while maintaining clarity.
For example: “A script reads: 'How can I help you today?' Record three versions: neutral customer service, warm and encouraging, and brisk/efficient. How would your delivery differ?”
4.Consistency across session boundaries
AI models need consistent voice characteristics across multiple recording days; this tests whether you can maintain consistency over time.
For example: “Your first session was 2 weeks ago. You're doing a second recording today. You notice your voice sounds slightly different. How would you verify you're consistent with Session 1?”
5.Script precision with natural delivery
TTS requires exact text matching, but natural speech can't sound robotic; this tests your balance between fidelity and naturalness.
For example: “A script has a sentence with awkward phrasing that doesn't feel natural to say. You could read it exactly as written or make it sound more conversational. What would you do?”
A task you may get
Record 30 seconds of Canadian French dialogue in two emotional registers (neutral and warm), ensuring clear audio quality and consistent voice characteristics.
How to prepare
- Test your recording setup in different times of day to identify and minimize background noise
- Collect 3-4 sample scripts in French and practice reading them with different emotional colorings
- Research the target AI application to understand what vocal characteristics the project values
The facts
- Pay
- $50/hr
- Hours
- Hourly
- Where
- Remote
- Open to
- CAN, CAN
- Field
- Language and Audio
- Posted
- 6/1/2026
- Places left
- 15
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.