$15–20/hr · Mercor · Hourly, 20 hours a week
An audiobook quality specialist who evaluates AI-generated Italian narration for speech quality and listener experience.
What you would do
- Listen to AI-synthesized Italian audiobooks with careful attention to narration quality
- Identify and document errors like mispronunciations, skipped words, and unnatural pacing
- Categorize issues and mark their exact location using provided annotation tools
- Assess overall satisfaction and pinpoint moments where narration feels artificial
- Report findings consistently across many hours of audio
Who they want
- Native or near-native Italian speaker with regular audiobook listening experience
- Working English proficiency for reporting and team communication
- Experience with annotation, transcription, proofreading, or similar detail work
- Patient ear with discipline to stay focused during long listening sessions
- Part-time availability of approximately 20 hours per week
Main skills
What the interview asks about
1.Detecting speech quality flaws
You must distinguish between genuine errors and acceptable variations to avoid flagging natural-sounding mistakes or missing real quality issues.
For example: “You're listening to a segment where the AI narrates a sentence with slightly rushed tempo. How do you decide whether this is a flaw worth documenting or a stylistic choice?”
2.Prioritizing listener experience
The role targets moments where narration breaks immersion, so you need judgment about what actually disrupts audiobook enjoyment.
For example: “You notice the AI consistently mispronounces a character's name in one chapter but corrects it later. Would you flag every instance or only certain ones, and why?”
3.Using annotation tools accurately
Clean data depends on precise categorization and timestamp marking so the AI training team can pinpoint and fix underlying issues.
For example: “You hear a skipped word at the 2:34 mark. Walk me through how you'd locate it, mark the exact moment, and categorize the error type.”
4.Sustaining accuracy over time
After 4+ hours of listening, fatigue can cause you to miss errors or become inconsistent in what you flag, compromising dataset quality.
For example: “It's hour 5 of your session and you're getting tired. You notice what sounds like a pronunciation issue but you're less certain than usual. How do you handle it?”
5.Assessing naturalness holistically
Beyond individual errors, you need to judge whether the overall narration sounds like a real human audiobook narrator or distinctly synthetic.
For example: “Across this 10-minute segment, individual words are pronounced correctly but the overall delivery sounds robotic. How do you document that as feedback?”
A task you may get
Listen to a 5-minute Italian audiobook excerpt and flag errors, categorizing each by type and providing approximate locations with brief descriptions of listener impact.
How to prepare
- Review 10-15 examples of common AI narration errors with timestamp markers to calibrate your ear
- Practice using the annotation tool interface with sample audio to build speed and accuracy
- Listen to high-quality and low-quality audiobook narration samples to build a reference for naturalness
- Reflect on your own audiobook listening habits and what details you naturally notice while enjoying a story
The facts
- Pay
- $15–20/hr
- Hours
- Hourly, 20 hours a week
- Where
- Remote · Remote — preferred: Italy
- Field
- Language and Audio
- Posted
- 9/14/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.