Training Turk

Biochemistry & Life Sciences Domain Expert

$65–105/hr · Mercor · Full time, 40 hours a week

A senior biochemistry or life sciences researcher improving frontier AI models' reasoning about real research work through quality assessment and benchmark design.

What you would do

  • Evaluate life sciences research scenarios and AI responses to identify weak reasoning, unsupported causal explanations, and implausible conclusions.
  • Write high-quality instruction specifications and golden solutions that define what correct life sciences reasoning looks like, grounded in how research is actually conducted.
  • Design challenging life sciences tasks and evaluation benchmarks that meaningfully test AI reasoning in your area of specialization.
  • Collaborate with client researchers and specialists to calibrate standards, resolve scientific ambiguities, and ensure consistent evaluation criteria.
  • Help build domain-specific evaluation tools and contribute to research team discussions about what frontier AI models should learn in your field.

Who they want

  • PhD in biochemistry, genetics, immunology, structural biology, neuroscience, or a closely related life science field.
  • 4+ years of substantive research at a recognized institution (university, medical center, national lab, or industrial research organization).
  • Senior level progression such as Senior Scientist, Staff Scientist, Research Scientist, Instructor, PI, or faculty with research direction ownership.
  • Record of peer-reviewed publications, strongly preferred including first-author work in strong-impact journals, demonstrating research rigor and credibility.
  • Deep specialization in a focused research area: protein structure, cellular mechanisms, immunology, genomics, neuroscience, or systems-level biology.

Main skills

Life sciences researchBiochemistry and molecular biologyPeer reviewed publication standards

What the interview asks about

  1. 1.Spotting pseudoscientific reasoning

    AI models generate plausible-sounding but scientifically wrong answers. Your job is to catch them. Interviewer assesses your ability to recognize mechanistic hand-waving and unsupported causal claims.

    For example: “An AI answer states: 'The hydrophobic core drives folding, so this mutation increases hydrophobicity and improves stability.' The reasoning is incomplete. What's missing and why would this fail peer review?”

  2. 2.Golden solution authoring

    Your golden solutions define what 'correct' means. Interviewer checks whether you can write rigorous, complete answers that reflect real research standards.

    For example: “Write a golden solution to: 'Explain the relationship between allosteric regulation and enzyme specificity in a specific enzyme system you know well.' What details would you include and exclude? What level of mechanistic explanation is appropriate?”

  3. 3.Benchmark design for scientific reasoning

    Good benchmarks require understanding what matters for scientific judgment. Interviewer assesses whether your benchmarks test meaningful aspects of expert reasoning.

    For example: “You're designing a benchmark to evaluate whether AI understands common experimental pitfalls in molecular biology. What specific scenarios would you include and why? How would you score an answer that misses a critical control?”

  4. 4.Experimental design assessment

    Much of life sciences is about whether an experiment is well-designed. Interviewer checks whether you can critique experimental logic rigorously.

    For example: “A proposed experiment uses CRISPR to knock out a gene and measures phenotype changes. The design lacks a control experiment. What would you recommend as a control and what specific issues does the current design fail to address?”

  5. 5.Domain depth and specialization

    You're hired for specialization, not generalism. Interviewer checks depth of knowledge in your specific area.

    For example: “Describe a recent published finding in your specialty that surprised you or challenged your prior assumptions. Why did it matter to your field and what experimental or conceptual innovation made it possible?”

  6. 6.Translating tacit expertise into criteria

    Much expert judgment is intuitive. Interviewer assesses whether you can make your judgment explicit and teachable.

    For example: “You instantly recognize weak reasoning in a biochemistry paper. Break down what signals tell you it's weak-is it methodology, logic, missing citations, lack of controls? How would you teach this pattern recognition to a junior scientist or to an AI system?”

A task you may get

Review a life sciences research scenario. Write a golden solution explaining reasoning rigor, correct approach, and scientific standards applied.

How to prepare

  • Review 2-3 recent first-author papers. Be ready to discuss experimental design choices, methodological decisions, and alternatives you rejected.
  • Identify a common misconception or hand-wavy explanation in your field. Prepare a clear explanation of why it's wrong and what correct mechanistic reasoning looks like.
  • Think of a benchmark or evaluation standard you've used in your research to judge quality or rigor. Describe how you would translate this into explicit, teachable criteria.
  • Prepare an example of a time you spotted a flaw in someone else's reasoning or experimental design that they had initially missed. What did you see and why did it matter?

The facts

Pay
$65–105/hr
Hours
Full time, 40 hours a week
Where
Hybrid · Bay Area, CA
Open to
USA
Field
Life, Physical, and Social Science
Posted
8/19/2026
Places left
10

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.