$65–105/hr · Mercor · Full time, 40 hours a week
A senior biochemistry or life sciences researcher improving frontier AI models' reasoning about real research work through quality assessment and benchmark design.
What you would do
- Evaluate life sciences research scenarios and AI responses to identify weak reasoning, unsupported causal explanations, and implausible conclusions.
- Write high-quality instruction specifications and golden solutions that define what correct life sciences reasoning looks like, grounded in how research is actually conducted.
- Design challenging life sciences tasks and evaluation benchmarks that meaningfully test AI reasoning in your area of specialization.
- Collaborate with client researchers and specialists to calibrate standards, resolve scientific ambiguities, and ensure consistent evaluation criteria.
- Help build domain-specific evaluation tools and contribute to research team discussions about what frontier AI models should learn in your field.
Who they want
- PhD in biochemistry, genetics, immunology, structural biology, neuroscience, or a closely related life science field.
- 4+ years of substantive research at a recognized institution (university, medical center, national lab, or industrial research organization).
- Senior level progression such as Senior Scientist, Staff Scientist, Research Scientist, Instructor, PI, or faculty with research direction ownership.
- Record of peer-reviewed publications, strongly preferred including first-author work in strong-impact journals, demonstrating research rigor and credibility.
- Deep specialization in a focused research area: protein structure, cellular mechanisms, immunology, genomics, neuroscience, or systems-level biology.
Main skills
What the interview asks about
1.Spotting pseudoscientific reasoning
AI models generate plausible-sounding but scientifically wrong answers. Your job is to catch them. Interviewer assesses your ability to recognize mechanistic hand-waving and unsupported causal claims.
For example: “An AI answer states: 'The hydrophobic core drives folding, so this mutation increases hydrophobicity and improves stability.' The reasoning is incomplete. What's missing and why would this fail peer review?”
2.Golden solution authoring
Your golden solutions define what 'correct' means. Interviewer checks whether you can write rigorous, complete answers that reflect real research standards.
For example: “Write a golden solution to: 'Explain the relationship between allosteric regulation and enzyme specificity in a specific enzyme system you know well.' What details would you include and exclude? What level of mechanistic explanation is appropriate?”
3.Benchmark design for scientific reasoning
Good benchmarks require understanding what matters for scientific judgment. Interviewer assesses whether your benchmarks test meaningful aspects of expert reasoning.
For example: “You're designing a benchmark to evaluate whether AI understands common experimental pitfalls in molecular biology. What specific scenarios would you include and why? How would you score an answer that misses a critical control?”
4.Experimental design assessment
Much of life sciences is about whether an experiment is well-designed. Interviewer checks whether you can critique experimental logic rigorously.
For example: “A proposed experiment uses CRISPR to knock out a gene and measures phenotype changes. The design lacks a control experiment. What would you recommend as a control and what specific issues does the current design fail to address?”
5.Domain depth and specialization
You're hired for specialization, not generalism. Interviewer checks depth of knowledge in your specific area.
For example: “Describe a recent published finding in your specialty that surprised you or challenged your prior assumptions. Why did it matter to your field and what experimental or conceptual innovation made it possible?”
6.Translating tacit expertise into criteria
Much expert judgment is intuitive. Interviewer assesses whether you can make your judgment explicit and teachable.
For example: “You instantly recognize weak reasoning in a biochemistry paper. Break down what signals tell you it's weak-is it methodology, logic, missing citations, lack of controls? How would you teach this pattern recognition to a junior scientist or to an AI system?”
A task you may get
Review a life sciences research scenario. Write a golden solution explaining reasoning rigor, correct approach, and scientific standards applied.
How to prepare
- Review 2-3 recent first-author papers. Be ready to discuss experimental design choices, methodological decisions, and alternatives you rejected.
- Identify a common misconception or hand-wavy explanation in your field. Prepare a clear explanation of why it's wrong and what correct mechanistic reasoning looks like.
- Think of a benchmark or evaluation standard you've used in your research to judge quality or rigor. Describe how you would translate this into explicit, teachable criteria.
- Prepare an example of a time you spotted a flaw in someone else's reasoning or experimental design that they had initially missed. What did you see and why did it matter?
The facts
- Pay
- $65–105/hr
- Hours
- Full time, 40 hours a week
- Where
- Hybrid · Bay Area, CA
- Open to
- USA
- Field
- Life, Physical, and Social Science
- Posted
- 8/19/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.