Training Turk

Mercor listing

Sonic Audit Specialist - Chinese

$22/hr

The role in one line

Evaluate transcription quality and audio alignment for machine learning training data in Mandarin Chinese.

Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.

What you would do

  • Listen to recorded speech and determine whether transcribers captured spoken content accurately
  • Verify that automated word-boundary detection placed segment start/stop points correctly
  • Tag errors using a fixed set of codes and document why each judgment passed or failed
  • Identify and correct misaligned segments through boundary adjustment or segment modification

Who they are looking for

  • Native or fluent Mandarin (mainland China) speaker with 5+ years living in China-dominant region
  • Strong written English proficiency to read rulebooks and compose clear feedback
  • Prior transcription, localization, linguistic annotation, or QA experience highly valued
  • Comfortable with 5-hour tasks requiring sustained focus and careful listening
  • Optional: phonetics background, audio editing tools familiarity, or AI annotation experience

What the interview is likely to probe

  1. 1.Distinguishing speech sounds

    Model training depends on correctly identifying when speakers use similar phonemes that could be confused, since training data with acoustic errors produces flawed models.

    Expect something like: “A speaker says what sounds like 'shi' but the transcriber wrote 'xi'. How would you listen to determine if this is an actual error or expected tonal/phoneme variation?”

  2. 2.Applying error codes consistently

    Researchers use your error tags to understand what mistakes appear most frequently, so inconsistent tagging undermines the dataset's utility for improving models.

    Expect something like: “Three separate tasks contain speakers using hesitation particles not captured in original transcriptions. How would you tag these identically across all instances?”

  3. 3.Identifying word boundaries in continuous speech

    Forced alignment data trains speech recognition to detect precisely when one word ends and the next begins, so incorrect boundaries cascade errors downstream.

    Expect something like: “In rapid speech where '我爱' and '瓦爱' sound nearly identical due to tone shifts, how would you verify the real word boundary and correct the automated segmentation?”

  4. 4.Writing technical justifications

    Your written explanations help researchers understand why a segment failed and what pattern to fix in retraining, turning judgment into actionable feedback.

    Expect something like: “A transcription shows '去' but the speaker likely said '居'. Explain in English why you'd reject this segment and what the training process should learn.”

Exercise you may get

Listen to a 2-minute Mandarin recording, review the provided transcription and word-level alignment, identify at least one transcription error and one boundary misalignment, tag them with error codes, and write brief justifications for each finding.

How to prepare

  • Review sample transcription and alignment errors to understand the specific error codes used
  • Practice distinguishing similar Mandarin phonemes and tones at varying speeds and noise levels
  • Study the rubric document that defines when variation counts as acceptable versus error
  • Complete calibration sessions comparing your judgments against expert reference answers

Facts

Pay
$22/hr
Commitment
hourly
Hours
10 per week
Work arrangement
remote · Remote
Posted
9/18/2026