$170/hr · Mercor · Hourly, 10 hours a week
You evaluate clinical documentation and AI-generated hospital notes for accuracy and completeness that meet hospitalist standards.
What you would do
- Review inpatient documentation including histories and physicals, daily progress notes, and discharge summaries for clinical accuracy and completeness
- Evaluate AI-generated inpatient records against the standard of what a practicing hospitalist would document in real patient care
- Annotate inpatient encounter data according to detailed guidelines, flagging omissions, hallucinations, and clinical inaccuracies
- Provide structured written feedback on AI documentation quality that engineering and product teams can use for system improvement
- Contribute expert input on annotation guidelines and help refine standards for ambiguous or edge-case scenarios
Who they want
- MD with hospitalist or internal medicine background and 2+ years of post-residency practice with active inpatient work
- US-based and able to work part-time remotely with minimum 10 hours per week availability
- Hands-on experience writing and reading H&Ps, progress notes, and discharge summaries as part of regular clinical work
- Advanced C1+ proficiency in English and one additional language: Spanish, Mandarin, Cantonese, Tagalog, Russian, Korean, Arabic, Portuguese, Polish, Japanese, or others listed
- Familiarity with Epic or similar EHR systems is preferred
Main skills
What the interview asks about
1.Clinical documentation accuracy assessment
AI notes can sound plausible but miss clinical details or reasoning that a hospitalist would include; distinguishing between good and false documentation is core to evaluation.
For example: “An AI progress note documents vital signs and medications correctly but omits the clinical assessment and plan for a patient with worsening renal function. What documentation gaps would you flag, and why would this note be clinically problematic?”
2.Hallucination and omission detection
AI systems can invent plausible-sounding clinical findings or miss actual findings documented in source materials. You must catch both overconfidence and incompleteness.
For example: “Review an AI-generated discharge summary that mentions a medication allergy never documented in prior notes and omits significant comorbidities from the history. How would you annotate these errors for system learning?”
3.Clinical reasoning and completeness
Complete documentation is not just a checklist; it reflects clinical reasoning about why findings matter and how they influence decisions.
For example: “An AI writes a complete history of present illness but the assessment section lists diagnoses without explaining how the clinical findings support them. How would you evaluate whether the AI understands clinical reasoning vs. just retrieving templates?”
4.Multilingual clinical judgment
Your clinical expertise applies across languages; you must recognize when translated or foreign-language documentation is clinically sound vs. when translation errors create medical risk.
For example: “Evaluate an AI-generated note in another language you are proficient in. How would you determine if terminology choices are clinically appropriate in that language, or if they represent errors in medical translation?”
5.Documentation standards and edge cases
Hospitalist documentation standards have subtle rules about completeness, conciseness, and what requires inclusion; you help define where AI systems should meet clinical standards.
For example: “An AI omits detailed neuro exam documentation because the patient had no neuro changes. Is this appropriate conciseness or a documentation gap that could create liability? How would you clarify this for AI systems?”
A task you may get
Review a sample AI-generated progress note for an inpatient scenario; identify documentation gaps, areas of incomplete clinical reasoning, and any hallucinations. Write structured feedback for an AI system.
How to prepare
- Select a recent progress note you wrote; analyze which clinical reasoning steps are explicit in documentation vs. implicit from your thought process. How would an AI learn these?
- Prepare examples of documentation omissions or shortcuts you have seen from less experienced providers; identify what clinical reasoning they miss.
- Gather 2-3 H&P or discharge summary elements that have tripped up AI systems; prepare to explain why the clinical nuance matters for patient care.
The facts
- Pay
- $170/hr
- Hours
- Hourly, 10 hours a week
- Where
- Remote
- Open to
- USA
- Field
- Medicine
- Project name
- Boron
- Posted
- 9/2/2026
- Places left
- 8
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.