Training Turk

Multilingual Primary Care Physician (MD) — Clinical Documentation & AI Evaluation

$170–190/hr · Mercor · Hourly, 10 hours a week

A practicing primary care physician reviews and annotates clinical documentation to evaluate how accurately healthcare AI systems capture real outpatient clinical work.

What you would do

  • Review outpatient clinical notes and AI-generated summaries for accuracy, completeness, and clinical soundness
  • Annotate encounter data against detailed clinical guidelines to identify errors, omissions, and inaccuracies
  • Identify hallucinations and clinically implausible statements in AI-generated documentation
  • Provide structured, written feedback to engineering teams on what needs improvement
  • Contribute expert input on documentation standards and help refine annotation guidelines

Who they want

  • MD with 2+ years experience since residency completion, actively practicing in outpatient family medicine, general internal medicine, or primary care
  • Fluency at C1+ level in English and one additional language including Spanish, Mandarin, Korean, Arabic, Portuguese, Cantonese, Tagalog, Russian, Haitian Creole, or Polish
  • Direct experience writing clinical notes in outpatient settings and working with EHR systems like Epic
  • Minimum 10 hours per week availability for remote work with flexible scheduling around your clinical practice
  • Preferred: 5+ years post-residency, note auditing experience, or prior AI/ML annotation work

Main skills

Clinical documentation reviewAI evaluationAnnotation guidelines

What the interview asks about

  1. 1.Hallucination and inaccuracy detection

    Healthcare AI systems generate text that sounds clinically plausible but may contain false medical facts; spotting these requires clinical expertise combined with careful attention to detail.

    For example: “An AI-generated note states a patient has 'history of myocardial infarction' but the actual chart shows only dyslipidemia and hypertension without any cardiac events. How would you categorize and document this error for the engineering team?”

  2. 2.Documentation standards and clinical gaps

    Outpatient documentation varies by specialty and practice setting; reviewers must distinguish between normal variation in documentation style and genuine clinical gaps or inaccuracies.

    For example: “An AI note on a hypertensive patient includes vital signs and medication list but omits adherence assessment. How would you distinguish this from normal documentation variation?”

  3. 3.Multilingual clinical judgment

    Medical terminology and clinical presentation descriptions differ across languages; evaluators must maintain clinical precision while assessing AI output in languages beyond English.

    For example: “You're annotating a Portuguese-language note describing a patient's symptoms as 'nervosismo com tonturas'. How would you evaluate whether the AI's clinical categorization aligns with how primary care providers in Portugal document similar presentations?”

  4. 4.Feedback for technical improvement

    Clinical reviewers must translate medical observations into patterns and actionable recommendations that engineers can use to improve model performance systematically.

    For example: “You find 15 AI notes where medication interactions are documented inconsistently. How would you frame this pattern as actionable feedback for the engineering team?”

  5. 5.EHR and clinical workflow knowledge

    AI documentation tools must work within real clinical systems; understanding how EHRs function and how clinicians actually document in them is essential to evaluating whether AI output is clinically usable.

    For example: “AI summaries include elements absent from your institution's Epic templates. How do you determine if this is a documentation gap or an Epic configuration issue?”

  6. 6.Clinical guideline interpretation

    Annotation guidelines themselves can be ambiguous or conflict with real clinical practice; experienced reviewers help clarify and refine standards based on actual outpatient work.

    For example: “Guidelines say flag 'incomplete chronic condition assessment,' but brief follow-ups for stable patients justify abbreviation. How do you propose clarifying this guideline?”

A task you may get

Review and annotate 4-5 deidentified outpatient encounters, evaluate AI-generated summaries against clinical standards, identify errors and omissions, and provide structured written feedback with specific, actionable suggestions.

How to prepare

  • Study your specialty's documentation standards and review how your current EHR handles clinical note templates and documentation workflows
  • Research documented cases of AI hallucinations in healthcare and understand how plausible-sounding but clinically false information appears in medical text
  • Collect 3-4 examples from your own practice of documentation gaps or quality issues you've identified, and practice describing them in technical terms
  • Review medical terminology and documentation conventions in the languages you'll be evaluated on, especially terminology that might differ from US practice

The facts

Pay
$170–190/hr
Hours
Hourly, 10 hours a week
Where
Remote
Open to
USA
Field
Medicine
Project name
Boron
Posted
9/2/2026
Places left
12

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.