Training Turk

Voice Actor: CX Agent Voice Cloning (Spanish - Latin America)

$50–100/hr · Mercor · Hourly

A Latin America-based female voice performer with native Latin American Spanish who records voice samples for text-to-speech AI model training.

What you would do

  • Record high-quality voice samples from scripts spanning conversational, narrative, instructional, and diverse tonal contexts
  • Deliver clear, expressive speech with precise control over tone, pacing, and pronunciation throughout recording sessions
  • Perform multiple script versions with deliberate variations in emotion, emphasis, and vocal style
  • Maintain neutral accent and consistent voice characteristics across multiple recording sessions

Who they want

  • Native speaker of Latin American Spanish with current residence in the region
  • Ability to deliver neutral, internationally intelligible Spanish without strong regional accent
  • Clear, natural, expressive voice with strong emotional range and diction control
  • Access to professional or near-professional recording setup with quality microphone, pop filter, and quiet environment
  • Availability for approximately 4-hour recording session, with possibility of additional similar session

What the interview asks about

  1. 1.Neutral accent maintenance

    TTS systems training on one accent variant need consistency; strong regional markers reduce intelligibility for diverse Spanish-speaking audiences.

    For example: “A script included expressions common in Argentine Spanish. How would you deliver that authentically while maintaining neutral pronunciation that's intelligible to Spanish speakers from Mexico, Colombia, and Peru?”

  2. 2.Emotional range with voice consistency

    AI models need the same voice across emotional variations; learners must distinguish emotion from voice identity.

    For example: “You recorded a customer service script in five emotional contexts: frustrated customer, calm support agent, confused caller, confident answer, and urgent situation. How did you vary emotional delivery while keeping your voice identifiable?”

  3. 3.Technical direction precision

    Recording for AI requires exact adherence to pacing and emphasis instructions; small deviations across takes compound training issues.

    For example: “Three versions of the same sentence required different emphasis: version A emphasize 'urgently', version B emphasize 'tomorrow', version C natural emphasis. How did you adjust without losing naturalness?”

  4. 4.Neutral delivery challenges

    Neutral accent is harder than natural speech; it requires conscious control to avoid regional patterns while remaining natural-sounding.

    For example: “You were recording in your native regional variant where certain vowel sounds differ from neutral Spanish. How did you adjust pronunciation while maintaining natural delivery and not sounding over-corrected?”

  5. 5.Audio quality problem-solving

    Recording quality determines dataset usability; performers must recognize and address audio issues without compromising consistency.

    For example: “Midway through a session, you noticed subtle background noise from air conditioning that hadn't been audible earlier. Did you stop and fix the environment, or continue? How did you decide?”

A task you may get

Record three versions of the same customer service script: neutral delivery, enthusiastic delivery, and empathetic delivery. Maintain consistent neutral accent and voice identity across all versions while varying emotional quality distinctly.

How to prepare

  • Practice recording yourself in your neutral Spanish variant versus your natural regional accent to identify differences and practice consistency
  • Review your recording setup: test microphone, check background noise, verify file format compatibility
  • Record practice clips and assess them for accent neutrality, clarity, and emotional variability while maintaining voice recognition
  • Study professional TTS voice examples to understand clarity, naturalness, and quality standards for model training

The facts

Pay
$50–100/hr
Hours
Hourly
Where
Remote
Open to
ARG, BOL, CHL, COL, CRI, CUB, DOM, ECU, SLV, GTM, HND, MEX, NIC, PAN, PRY, PER, URY, VEN
Field
Language and Audio
Posted
6/1/2026
Places left
10

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.