Training Turk

Mercor listing

General Clinician (MD/DO) - HLS Expert Pool

$150/hr

The role in one line

Licensed physician evaluating how well AI systems perform on clinical tasks through assessment design, dialogue review, and reasoning annotation.

Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.

What you would do

  • Break clinical problems into measurable grading criteria that capture what makes an answer correct or unsafe
  • Review AI conversations and outputs, flagging inaccuracies, missed diagnoses, and clinically dangerous statements
  • Document your own clinical thinking on cases, showing differential diagnosis and reasoning steps for AI training
  • Write specialty guidelines and define edge cases so annotation teams stay consistent
  • Design complex cases that test model performance on uncommon or tricky presentations

Who they are looking for

  • MD or DO with completed residency in any specialty and active unrestricted medical license
  • Minimum 3 years post-residency clinical experience, actively practicing or recently practicing
  • Able to write clear clinical reasoning that non-physician reviewers can follow and verify
  • Fluent in written and spoken English with availability for 20+ hours weekly
  • Board certification, US licensure, primary care or emergency medicine background, or prior medical writing experience preferred

Skills this role asks for

clinical judgmentmedical ai evaluationannotation frameworksclinical reasoning documentationdialogue assessmentsafety reviewguideline development

What the interview is likely to probe

  1. 1.Creating assessment criteria from clinical complexity

    This measures whether you can translate fuzzy clinical judgment into objective, verifiable standards that remain reliable when many annotators use them.

    Expect something like: “A diabetic patient presents with chest pain and shortness of breath. What checkpoints would you use to grade whether an AI system's response captures all dangerous possibilities without requiring an emergency workup for every visit?”

  2. 2.Identifying dangerous AI errors and overconfidence

    Most clinical AI failures aren't exotic - they're ordinary questions answered with misplaced certainty. Your job is catching that with your actual clinical experience of what goes wrong.

    Expect something like: “An AI system correctly identifies bacterial pneumonia and recommends appropriate antibiotics, but doesn't mention that you should verify the patient isn't pregnant before prescribing that class. How would you flag this as a safety issue in your review?”

  3. 3.Documenting clinical reasoning for machine learning

    AI learns from how clinicians actually think through cases - not just the diagnosis, but the differentials you ruled out and why. Your annotations teach it to reason like you do.

    Expect something like: “A 28-year-old woman reports fatigue, weight loss, and fever for three weeks. Walk me through what differentials you'd consider first, which you'd deprioritize for her age and presentation, and what would push you toward or away from each diagnosis.”

  4. 4.Setting specialty standards and edge cases

    Guidelines you author keep a large team of clinicians annotating consistently. Gaps in your edge-case definitions create inconsistency that breaks the signal.

    Expect something like: “You're writing assessment criteria for depression screening. Should the standard be the same for an 18-year-old college student, a 45-year-old with untreated diabetes, and a 72-year-old with cognitive decline? What changes?”

  5. 5.Designing cases that probe model limits

    Simple cases don't expose AI reasoning. You need cases that are genuinely tricky, uncommon, or have competing interpretations - the cases where models are most likely to fail.

    Expect something like: “You're designing cases to test emergency medicine reasoning. What would be a realistic but unusual presentation that forces the model to consider rare diagnoses, distinguish between similar-looking conditions, or recognize when escalation is needed?”

Exercise you may get

Review an AI response to a chest pain case. The system recommends stress testing and reassurance. Write grading criteria to evaluate if the response is complete and safe, identifying what assumptions might be dangerous.

How to prepare

  • Review several published medical education assessments or board exam questions in your specialty to see how standards define correct answers
  • Prepare examples of cases where AI systems might sound confident but miss a critical diagnosis or safety consideration in your field
  • Refresh your knowledge of current clinical guidelines and standards of care in your specialty, including any recent guideline updates
  • Write a brief sample of your own clinical reasoning on a case - show how you'd think through differentials and next steps for someone unfamiliar with medicine

Facts

Pay
$150/hr
Commitment
hourly
Hours
20 per week
Work arrangement
remote · Remote — worldwide
Eligible locations
USA, USA
Posted
9/17/2026