Training Turk

Voice Actor: CX Agent Voice Cloning (Swiss German)

$50–150/hr · Mercor · Hourly

Native Swiss German speaker recording speech samples for training AI text-to-speech models serving customer-experience applications.

What you would do

  • Record voice samples spanning conversational, narrative, and instructional scenarios in Swiss German
  • Deliver articulate and emotionally nuanced speech with consistent regional dialect and pronunciation
  • Maintain vocal consistency across multiple session hours and re-recording iterations
  • Follow technical audio recording guidelines for equipment, environment, and output formatting
  • Execute multiple takes varying emotion, emphasis, and pacing as requested by researchers

Who they want

  • Native Swiss German speaker currently living in Switzerland
  • Natural and expressive voice quality with clear articulation (prior voice or broadcast experience valued but not required)
  • Access to semi-professional recording setup featuring quality microphone, sound-isolated space, and basic equipment
  • Mastery of intonation, diction, and emotional coloration suited to customer-service contexts
  • Flexibility for an initial 4-hour session with potential for follow-up recording work of similar duration

Main skills

Swiss german dialectVoice synthesis trainingAI speech generation

What the interview asks about

  1. 1.Zurich dialect and pronunciation precision

    AI learns regional phonology from data; your reliable production of distinctly Swiss German features - schwa sounds, consonant clusters, rhythm patterns - is what makes synthesis convincing to speakers.

    For example: “How do you differentiate your pronunciation of 'grueezi' from Standard German 'gruesse'? Walk through the vowel and consonant choices that mark your speech as distinctly Swiss.”

  2. 2.Shifting vocal tone for different scenarios

    Customer-service speech spans greeting warmth, technical patience, and problem-solving urgency; models learn these tonal shifts from your performance.

    For example: “Here's a script for a customer reporting an issue. First deliver it as a warm greeting, then as a focused troubleshooter. What vocal choices differ between the two?”

  3. 3.Audio recording execution and troubleshooting

    Clean recordings directly affect model quality; you must manage your microphone position, room noise, and technical settings to produce usable training data.

    For example: “You're recording at home with a USB microphone. Describe your setup strategy - how do you minimize echo, reduce background hum, and position yourself consistently for every take?”

  4. 4.Naturalness versus script fidelity

    Robotic word-by-word delivery doesn't train good synthesis; AI learns natural pacing from how you handle scripts - where you breathe, what you emphasize, rhythm choices.

    For example: “You have a 45-second customer-service script. Mark where you'd naturally pause, which phrases you'd stress, and describe your pacing strategy to sound spontaneous not recited.”

  5. 5.Voice cloning technology and consent

    Accepting voice cloning for this AI agent demonstrates informed consent and grasp of how your vocal characteristics become the model's output voice.

    For example: “Describe what you understand voice cloning to mean for this project. How will your recordings be transformed into synthetic speech, and what are your guarantees about use?”

A task you may get

Record a brief customer-service interaction (caller inquires about a billing issue, you provide calm, clear guidance) in Swiss German. Deliver the same scenario twice: first as a formal business tone, second as a friendly but professional colleague.

How to prepare

  • Set up and test your home recording equipment - microphone, audio interface, recording software - to achieve professional-quality output
  • Study customer-service interactions in Swiss German to internalize authentic dialect and conversational pacing in that context
  • Practice delivering varied emotional tones in Swiss German without code-switching to Standard German
  • Understand voice cloning basics and how your voice data contributes to AI speech synthesis

The facts

Pay
$50–150/hr
Hours
Hourly
Where
Remote
Open to
CHE
Field
Language and Audio
Posted
7/30/2026
Places left
10

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.