$20–25/hr · Mercor · Hourly, 20 hours a week
Audio QA specialist evaluating AI-narrated text-to-speech audiobooks for synthesis errors and listener experience.
What you would do
- Listen to AI-narrated audiobooks and assess how naturally and accurately the narration sounds
- Identify synthesis errors including mispronounced words, skipped or inserted content, incorrect abbreviation expansion, unnatural phrasing and intonation
- Mark error locations precisely using annotation tools and categorize each issue type
- Rate overall listener satisfaction and indicate where narration flows smoothly versus where it disrupts the listening experience
- Maintain attention to detail and accuracy consistently across many hours of audio content
Who they want
- Native or near-native English fluency in US/North American English with genuine audiobook listening habits
- Experience with annotation, transcription, proofreading, linguistics or similar detail-oriented review work
- Sharp ear and discipline to stay accurate through long listening sessions without losing focus
- Reliable availability for approximately 20 hours per week on flexible part-time schedule
- Clear written English for reporting findings and communicating with the team
Main skills
What the interview asks about
1.Detecting synthesis errors in context
Ability to spot subtle errors in flowing speech requires both linguistic knowledge and consistent attention, critical to training better synthesis models.
For example: “You're listening to a chapter where the narrator should say 'the patient received a 500 mg dose'. You hear '500 em-jee dose'. Is this an error? How would you categorize it and mark the exact span?”
2.Distinguishing natural from unnatural narration
Assessing listener satisfaction requires experience with how real audiobooks should sound, differentiating minor errors from experience-breaking moments.
For example: “The AI reads 'He said, "I'll be there at seven p.m."' with strange stress on the 'p' sound in p.m. versus when it reads the same abbreviation naturally in a different sentence. What makes one version jarring and the other acceptable?”
3.Categorizing error types accurately
Consistent error classification helps downstream teams understand failure patterns and target synthesis improvements.
For example: “In a medical chapter, you hear 'myocardial infarction' pronounced as 'myo-car-dee-ul infraction'. Is this a mispronunciation, a phonetic substitution, or both? How would you mark and describe it?”
4.Maintaining accuracy across extended sessions
QA work loses value if attention lapses cause missed errors or miscategorization after several hours of listening.
For example: “You're 6 hours into an 8-hour evaluation shift. You notice your error-detection rate feels slower. What strategies keep your assessment quality consistent through the final hours?”
A task you may get
Listen to 30 minutes of AI-narrated audiobook content (provided), identify and mark at least 10 errors using categorization labels, rate overall listening experience, and write a brief summary of where synthesis quality was strongest and weakest.
How to prepare
- Listen to 2-3 professionally narrated audiobooks and one AI-narrated sample to calibrate your ear to quality standards
- Familiarize yourself with common text-to-speech error patterns: abbreviations, numbers, contractions, proper nouns
- Practice using a simple annotation tool to mark and categorize content issues with precise time stamps
- Reflect on your own audiobook listening: what makes narration feel natural versus when you notice flaws
The facts
- Pay
- $20–25/hr
- Hours
- Hourly, 20 hours a week
- Where
- Remote · Remote — preferred: United States
- Field
- Language and Audio
- Posted
- 9/14/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.