$100–150/hr · Mercor · Full time, 40 hours a week
You guide an AI lab's research into legal reasoning by designing tasks, reviewing model outputs for real-world correctness, and building benchmarks that teach models to reason like a specialist lawyer.
What you would do
- Review legal task quality and model outputs, identifying reasoning gaps that surface-level reading obscures
- Write clear task specifications and golden solutions reflecting real legal work
- Design challenging evaluation benchmarks and legal-specific assessment sets with research teams
- Translate professional legal judgment into explicit, learnable criteria
- Work with researchers to develop and calibrate legal-specific tools and evaluation methodologies
Who they want
- Minimum five years of professional legal practice following bar admission at a firm, corporation, government, or court
- Genuine specialization in one practice domain such as corporate transactions, litigation, regulatory compliance, intellectual property, employment law, or tax
- Clear career progression to senior level with demonstrated matter ownership and mentorship experience
- Juris Doctor from accredited law school and active bar admission in at least one U.S. state
- Working familiarity with large language models and experience assessing legal reasoning quality
What the interview asks about
1.Defining correctness in specialized law
Models excel at pattern matching but struggle with deep reasoning; you must make professional judgment explicit enough for a researcher to measure and improve.
For example: “M&A opinion identifies regulatory hurdles but circular on risk-benefit analysis. Score this. Write golden solution teaching step-by-step statutory framework reasoning.”
2.Task specification for legal reasoning
Vague or ambiguous task specs lead models and evaluators astray; your specs must be crisp enough for consistent scoring yet realistic enough to matter in practice.
For example: “You're designing a benchmark task on securities compliance. Walk me through how you'd write the prompt, the golden solution, and the scoring rubric such that two independent reviewers score the same model output identically 95% of the time.”
3.Spotting plausible-sounding error
The hardest failures are those that read well on the surface; catching them requires deep domain judgment.
For example: “Model gives competent precedent summary but misses critical distinction changing outcome. Cites right cases; holdings subtly off. How catch in QA? What flag?”
4.Calibrating across specialties
You're working with researchers in multiple legal domains; inconsistent standards create confusing data and slow model improvement.
For example: “Build benchmarks for employment and tax law. Evidence standards differ. Document calibration so researcher between domains knows which standards apply.”
5.Career progression and matter ownership
You're advising on deep legal work; the interviewer needs confidence you've owned complex matters, not just assisted on them.
For example: “Tell me about the most complex legal matter or transaction you've owned from start to finish: what was the core legal question, what arguments did you develop, and what outcome did you achieve?”
A task you may get
Write task specification and golden solution for realistic legal scenario: client situation, legal question, analysis framework, key precedent, counterarguments, conclusion. Format for objective model scoring.
How to prepare
- Prepare a 10-minute walk-through of your most complex legal matter or transaction, emphasizing the judgment calls and reasoning path you took
- Read 2-3 recent decisions in your specialty domain and identify where a model trained on general legal text would likely go astray
- Think through a real client scenario where surface-level legal analysis would produce the wrong outcome; prepare to explain what the model would miss
- Reflect on how you'd explain a subtle area of your specialty to a research team that knows no law: what mental model or framework would you use?
The facts
- Pay
- $100–150/hr
- Hours
- Full time, 40 hours a week
- Where
- Hybrid · Bay Area, CA
- Open to
- USA
- Field
- Law
- Posted
- 8/12/2026
- Places left
- 5
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.