$60–80/hr · Mercor · Part time
Physical scientist who evaluates AI model reasoning and creates training tasks based on experimental expertise.
What you would do
- Review model outputs for scientific accuracy, methodological soundness, and reasoning quality
- Design scenario-based tasks grounded in realistic experimental or theoretical work to test model understanding
- Provide technical feedback identifying reasoning gaps and scientific errors in model performance
- Contribute task specifications and golden solutions that define correct scientific reasoning
- Assess how well the model grasps measurement practices, uncertainty, and experimental design principles
Who they want
- Professional research experience in experimental methods, quantitative analysis, and scientific reporting in physics, chemistry, materials science, or similar domain
- Doctoral degree (PhD or equivalent) in a physical science discipline strongly preferred
- Ability to communicate technical findings clearly in written form for research teams who may lack your specific domain background
- Comfort working independently and remotely, managing your own schedule across rolling projects with 15-30 hours per week typical commitment
- Strong judgment for distinguishing rigorous scientific reasoning from fluent-sounding but technically unsound explanations
Main skills
What the interview asks about
1.Assessing model physics understanding
You need to recognize when a model generates coherent text about a physical system but misses material factors or uses flawed logic a real researcher would spot. This tests your ability to evaluate depth of reasoning.
For example: “A model discusses electron scattering and phonon behavior in composites but omits interface resistance and misapplies single-component assumptions. What gaps would you flag and what test case would prove understanding?”
2.Designing rigorous assessment scenarios
Generic textbook problems don't test model reasoning at the depth a scientist uses daily. You must design tasks that capture real complexity and ambiguity from your own research experience.
For example: “Design a task to evaluate AI understanding of measurement uncertainty in spectroscopy. What limitations, noise, and systematic errors would you include? How do you distinguish superficial awareness from genuine mastery?”
3.Identifying reasoning versus pattern-matching
AI models excel at statistical patterns but may lack causal understanding. You need to spot when an answer looks polished but commits category errors or misses crucial physics.
For example: “A model discusses pressure-induced transitions but misses kinetic barriers and equilibrium assumptions. How would you structure feedback to reveal true understanding of transition mechanisms?”
4.Translating science into evaluation criteria
Your judgment is tacit and hard-won from years in the lab. You need to make it explicit enough that others can apply consistent standards across many evaluations.
For example: “You're expert in X-ray diffraction interpretation. Write an evaluation rubric for non-specialists to assess XRD model responses, covering both quantitative and qualitative reasoning that matters most.”
5.Remote independent project management
You'll juggle multiple concurrent assessments without direct oversight, so you need strong self-direction and clear communication about progress and blockers.
For example: “You're assigned three models to evaluate in four weeks, but one project is harder than expected, pushing your timeline. How would you communicate the issue, what help would you need, and what adjustments would you propose?”
6.Experimental design methodology
Models often treat experimental design as a checklist rather than reasoning about controls, confounds, and statistical power. You must evaluate this deeply.
For example: “A model proposes a catalyst experiment with sample sizes and measurement techniques but ignores confounding factors and isolation methods. What experimental rigor issues would you flag and what scenario tests understanding of proper controls?”
A task you may get
Review three AI responses on physics or chemistry (thermodynamics, procedures, materials). Identify what's correct, what reasoning gaps exist, and design a task to reveal each gap.
How to prepare
- Gather 3-4 research problems from your own work that required careful reasoning about methodology, uncertainty, or confounding factors
- Review published examples of AI model outputs on science topics and practice articulating what distinguishes sound reasoning from sophisticated-sounding but flawed answers
- Prepare a brief summary of your core research area, the types of questions you worked on, and the methodological challenges that matter most
- Practice explaining technical concepts to ML researchers without domain knowledge, since you'll need to communicate feedback clearly
The facts
- Pay
- $60–80/hr
- Hours
- Part time
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Role type
- Talent network
- Posted
- 2/27/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.