$70–100/hr · Mercor · Part time, 40 hours a week
Physicist with graduate training and active research record creating and evaluating research-level physics problems for AI model assessment.
What you would do
- Write original physics problems from your expertise area, including system description, assumptions, regime, and a reference solution for AI grading
- Grade model-written derivations and numerical work, assessing whether approximations are valid in the claimed regime
- Check limiting cases and unit consistency to catch errors models might hide in correct final expressions
- Identify when models reach right answers through invalid reasoning that breaks under changed assumptions
- Document your analysis and reasoning clearly so others can verify and learn from your work
Who they want
- Graduate degree in physics or closely related field, with hands-on research experience
- Recent and verifiable publication record; papers listed will be checked against public record
- Command of the specific methods and phenomena you're evaluating, not just adjacent topics
- Working proficiency with LaTeX, Python, SymPy, and Jupyter; some VS Code familiarity helpful
- B2-level English proficiency including clear written reasoning and technical explanation
Main skills
What the interview asks about
1.Detecting flawed reasoning
Models can produce correct final expressions through invalid derivations; identifying this requires deep understanding of when each step actually holds.
For example: “A model solved a quantum mechanics problem and got the right energy eigenvalue, but used an approximation that's only valid for weak coupling, not the strong-coupling regime the problem specified. How would you have caught this in grading?”
2.Constructing research-grade problems
Problems must test genuine research reasoning, not textbook pattern-matching; your problem-writing shows whether you understand what makes a challenge actually hard.
For example: “Write a research-level problem from your subfield that tests one specific technique or insight that frontier models might struggle with. Walk me through your problem statement and reference solution.”
3.Verifying dimensional analysis
Unit errors propagate invisibly through complex derivations; spotting impossible scaling or dimensionless results masquerading as dimensional quantities catches subtle mistakes.
For example: “A model derived a formula for energy density that, when checked dimensionally, involved a mass-to-length ratio. What would that error suggest about the derivation, and how would you have flagged it?”
4.Assumption validity across regimes
Physics breaks in different ways depending on regime; checking whether an approximation survives regime changes reveals whether reasoning is robust or accidentally correct.
For example: “In your field, you've published on systems where an approximation holds for small fields but fails for strong fields. Given a model's solution, how would you test it across both regimes?”
A task you may get
Write a research-level physics problem from your publication area with full problem statement, assumptions, regime description, and a complete reference solution (800-1200 words), then explain what makes this problem genuinely challenging for AI.
How to prepare
- Prepare your 3-5 most representative papers with arXiv IDs or DOIs, ready to cite as examples of your methods
- Review a recent physics preprint outside your core expertise to practice reading derivations critically
- Compile a list of 2-3 common errors you catch when reviewing papers in your field
The facts
- Pay
- $70–100/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.