Training Turk

Audio Engineer (Speech / TTS Audio Specialist) – French

$50/hr · Mercor · Hourly

You edit and quality-check French voice recordings for machine learning speech systems, ensuring technical precision and linguistic authenticity at scale.

What you would do

  • Remove silence, breath noise, clicks, distortion, and background artifacts from French speech files using restoration software
  • Conduct detailed quality checks for clipping, loudness consistency, background noise, and technical defects
  • Evaluate whether recordings demonstrate accurate French pronunciation, natural intonation, and native-like fluency
  • Deliver processed files meeting strict technical specifications for ML training systems
  • Process high-volume batches of recordings while maintaining consistent quality across deliverables

Who they want

  • Fluent native or near-native French speaker with refined listening skills for assessing French speech characteristics
  • Hands-on experience editing voice or speech recordings using professional audio restoration tools such as iZotope RX, Pro Tools, or equivalent software
  • Understanding of common audio defects: clipping, distortion, background noise, inconsistent levels in dialogue recordings
  • Experience working with large audio datasets or structured, high-volume workflows
  • Familiarity with TTS systems, speech ML projects, or professional audio dataset creation workflows

Main skills

French language audio qualitySpeech editing and cleanupAudio restoration techniques

What the interview asks about

  1. 1.French speech quality judgment

    You need to hear what is authentically French versus what sounds unnatural or accented, which determines whether recordings train models correctly on native speech patterns.

    For example: “You receive 30 recordings of the same French phrase spoken by different voices. One has uneven nasalization, another has rushed pacing. How would you assess whether each is acceptable for a speech model?”

  2. 2.Audio artifact prioritization

    Some defects like clipping permanently destroy signal and must be removed; others like subtle breath noise can remain if they don't violate spec. You must triage repair effort.

    For example: “A recording has slight background hum throughout and one section with obvious breath intake. Which do you address first and why?”

  3. 3.Throughput and precision balance

    You're expected to complete dozens or hundreds of files, but quality slip-ups compound across a training dataset. You must maintain speed without sacrificing accuracy.

    For example: “Your project requires processing 200 files in two weeks. After the first 30, you notice the background noise in some files near target volume. Do you re-process them, adjust your workflow, or continue?”

  4. 4.Technical specification adherence

    ML pipelines require consistent specifications: bit depth, sample rate, loudness range. Deviations break downstream processes even if they sound okay to your ear.

    For example: “The project spec requires loudness between -23 and -20 LUFS with no peaks above -1dB. One file hits -1.2dB at a breath sound. Is trimming the breath enough or do you need another approach?”

A task you may get

Edit a raw French speech recording that contains breath noise, background hum, level inconsistency, and one clipped peak. Demonstrate your cleanup decisions and explain how you balanced removing defects against preserving natural French speech character.

How to prepare

  • Listen to 3-5 samples of professional French speech recordings at different quality levels, noting what makes one sound authentic versus produced
  • Review technical specs for a TTS dataset (loudness, bit depth, sample rate) and practice mixing a file to meet those exact tolerances
  • Spend 30 minutes in iZotope RX or your preferred tool practicing de-clicking, de-breathing, and noise reduction on a sample French speech file
  • Identify 2-3 French pronunciation features (nasal vowels, liaison, final consonant behavior) and listen for how they appear or vanish in various recordings

The facts

Pay
$50/hr
Hours
Hourly
Where
Remote
Field
Language and Audio
Posted
6/1/2026
Places left
5

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.