$50–100/hr · Mercor · Hourly
A native Peruvian Spanish voice actor who records high-quality training speech for text-to-speech and conversational AI systems.
What you would do
- Record voice samples from diverse scripts in natural, expressive Peruvian Spanish
- Perform multiple takes with emotional variation and stylistic changes
- Maintain consistent accent, tone, and delivery quality across the recording session
- Follow technical recording guidelines for microphone setup and file formatting
- Deliver instructional, conversational, and narrative content with appropriate vocal choices
Who they want
- Native female Peruvian Spanish speaker currently based in Peru
- Clear, natural voice with strong emotional range and expressiveness
- Professional or near-professional recording setup including quality microphone and quiet environment
- Precise control over intonation, diction, pacing, and pronunciation
- Availability for approximately 4-hour recording sessions with possibility of future sessions
Main skills
What the interview asks about
1.Accent consistency and authenticity
AI voice cloning models must train on consistent accent; drift or unnatural phonetics degrade the cloned voice quality.
For example: “You're recording a 4-hour session with 200 sentences in varied Peruvian Spanish dialects and regional vocabulary. How would you maintain consistent accent and pronunciation across fatigue and topic changes?”
2.Emotional delivery and vocal range
Conversational AI agents need diverse vocal styles; limited emotional range makes cloned voices sound robotic or unnatural.
For example: “A script reads 'I understand your frustration' and you need 5 takes: one warm and empathetic, one neutral, one slightly frustrated, one formal, one energetic. How would you approach vocal variation while keeping the same words clear?”
3.Recording environment and mic technique
Poor recording quality wastes studio time; background noise, plosives, or clipping make training data unusable.
For example: “Your recording setup has a quality microphone in a small bedroom. What acoustic challenges would you anticipate, and what practical steps would you take to minimize background noise and plosives?”
4.Consistency across multiple sessions
If the AI lab schedules a second session later, vocal consistency ensures models train on coherent data rather than inconsistent voices.
For example: “You record session one and three weeks later are called back for session two. What would you document from session one to help you match your voice, accent, and delivery style?”
5.Script variety and register switching
TTS models need training across conversational, instructional, and formal registers; poor switching makes cloned voices sound unnatural.
For example: “In one session you record customer service responses, technical documentation, and friendly chitchat. How do you maintain natural delivery while switching registers without sounding like different speakers?”
A task you may get
Record 20-30 short sentences in Peruvian Spanish using provided scripts covering conversational and instructional styles, perform 3 takes per sentence with different emotional tones, and verify audio quality meets professional standards.
How to prepare
- Test your recording environment for background noise and acoustic issues
- Practice delivering the same sentences with 3-4 different emotional tones naturally
- Review samples of professional TTS voice recordings to understand quality and consistency standards
- Familiarize yourself with audio file formats and naming conventions required by AI training systems
The facts
- Pay
- $50–100/hr
- Hours
- Hourly
- Where
- Remote
- Open to
- PER
- Field
- Language and Audio
- Posted
- 6/21/2026
- Places left
- 3
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.