Training Turk

Legal Domain Expert — AI Training & Evaluation

$60–100/hr · Mercor · Full time, 40 hours a week

You evaluate and improve how AI models reason about legal work, writing benchmarks and instruction standards for a frontier AI lab.

What you would do

  • Review legal knowledge work tasks and model outputs for correctness, spotting reasoning gaps and answers that only appear sound
  • Write instruction specifications and develop golden solutions that exemplify strong legal analysis across specific practice areas
  • Design evaluation sets and benchmarks that measure whether models genuinely understand legal reasoning
  • Collaborate with AI researchers and domain specialists to translate legal judgment into measurable, teachable criteria

Who they want

  • JD from a top law school plus 5+ years working in substantive legal roles at a law firm, corporate counsel, or government agency
  • Genuine specialization in at least one practice area: corporate, litigation, regulatory, IP, employment, or tax
  • Senior-level progression: partner, counsel, senior associate, or general counsel with clear ownership of complex matters
  • Active U.S. state bar admission in good standing and hands-on professional use of large language models
  • Bay Area-based with ability to work on-site multiple days weekly; excellent written communication skills

Main skills

Legal domain expertiseAI evaluation and benchmarkingInstruction specification writing

What the interview asks about

  1. 1.Legal reasoning quality assessment

    You must distinguish between answers that sound plausible and answers that would survive professional scrutiny, which is core to evaluating whether an AI model actually grasps legal reasoning or is just fluent.

    For example: “An AI model drafted a memo analyzing successor liability in an asset purchase. The analysis was well-structured and cited relevant case law. What specific gaps in reasoning would make you flag this as not ready for a junior associate to present to a client?”

  2. 2.Instruction specification and golden solution development

    The quality of your instruction specs directly determines whether researchers can build accurate benchmarks and whether models improve correctly; vague specs fail both.

    For example: “You're writing an instruction spec for a regulatory compliance task. A client faces potential SEC disclosure obligations. How do you structure the task so the benchmark scores both junior lawyers and AI models correctly?”

  3. 3.Specialization depth and domain judgment

    You're hired for deep practice expertise, not generalist legal knowledge. The benchmarks you design must reflect real work complexity in your specialty or they won't evaluate models on what matters.

    For example: “The research team asks you to expand benchmarks into a practice area adjacent to your specialty but where you have minimal experience. How do you respond, and what would you recommend instead?”

  4. 4.Translation of tacit judgment into teachable criteria

    Researchers and engineers need explicit, measurable criteria. If you can only say I know it when I see it, your benchmarks won't improve model performance.

    For example: “A model's answer on a corporate tax issue shows good reasoning by some measures but you'd reject it as junior counsel. What three explicit criteria would you document so the research team can measure this gap?”

  5. 5.Cross-functional collaboration with AI teams

    Researchers and engineers need to understand your domain constraints and you need to understand their technical limitations; misalignment wastes weeks of work and produces unusable benchmarks.

    For example: “A researcher proposes a benchmark based on a recent Supreme Court ruling that changed your specialty overnight. How do you work with them to scope what's feasible to evaluate in an AI model vs. what requires live legal judgment?”

A task you may get

Given a sample legal task in your specialty, write a 200-word instruction spec and a golden solution; identify two scenarios where an AI model might produce a plausible-sounding but legally flawed answer.

How to prepare

  • Prepare a detailed example from your practice of work that was harder than it initially looked, and where flawed reasoning could have cost a client money or legal exposure.
  • Gather 3-5 recent cases or regulatory developments in your specialty; be ready to explain why this context matters for evaluating AI models.
  • Draft a short note on the hardest part of judging legal quality: what makes you confident in your own analysis, and how would you teach someone else to see that confidence?

The facts

Pay
$60–100/hr
Hours
Full time, 40 hours a week
Where
Hybrid · Bay Area, CA
Open to
USA
Field
Law
Posted
8/11/2026
Places left
9

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.