$20–160/hr · Mercor · Task based, 15 hours a week
Assess health-related AI conversations from a patient perspective, judging clarity, honesty, and usefulness.
What you would do
- Rate whether AI responses actually answered the question posed, using language a non-expert could follow
- Detect where AI hedged, retreated, or offered reassurance instead of straight answers when users challenged it
- Evaluate tone: whether the exchange left someone calmer and prepared to act, or anxious and confused
- Examine what the response deliberately omitted or sidestepped
- Write specific, quotable critiques explaining what a better response would have said
Who they want
- Strong English writing skills; ability to explain judgments in 2-4 sentences rather than one-liners
- Careful reader who notices gaps and omissions, not just obvious errors
- Experience applying rubrics, grading standards, or quality-review processes consistently
- Reliable availability to focus intensely on batches of work over defined timeframes
- Preferably: background teaching, editing, writing, support, or patient advocacy work
Main skills
What the interview asks about
1.Detecting evasion and hedging
Clinicians miss this because they read for accuracy; you catch where the AI sounds helpful but actually avoids the hard answer.
For example: “User asks if a symptom is serious. AI responds warmly but vaguely about checking with doctor. What warning signs should be explicit?”
2.Clarity for non-specialists
Medical jargon and hedging statements sound responsible but leave people no clearer than before.
For example: “A response says: "This might indicate a need for investigation into your metabolic function." A patient reading that has no idea whether they should see a doctor, test at home, or wait. How do you score this for clarity, and what concrete language would help?”
3.Tone and emotional outcome
Conversations that sound supportive but leave someone worse off psychologically harm more than silence.
For example: “A user describes anxiety about a symptom. The AI validates their feeling but then lists 12 possible causes without prioritizing which are rare vs. common, or which need urgent attention. The user leaves more anxious. How would you document this failure?”
4.Written justification depth
One-line verdicts don't help teams fix the problem; detailed explanation shows you understood what went wrong.
For example: “You rated a response as 'partially helpful' rather than 'excellent.' You must explain why. What specific phrase or omission drove your rating, and what would fix it in 2-3 sentences?”
A task you may get
Rate 5-10 health conversations on clarity and directness; write justifications for any rating below excellent.
How to prepare
- Practice reading health-related text critically and taking detailed notes on what was said vs. what was omitted
- Work through 2-3 practice conversations applying a sample rubric and writing justifications under time constraints
- Study how health information for laypeople should be framed: plain language, specific next steps, clear risk levels
- Review examples of evasive vs. direct communication and practice spotting the difference
The facts
- Pay
- $20–160/hr
- Hours
- Task based, 15 hours a week
- Where
- Remote · Remote — worldwide
- Posted
- 9/11/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.