$80–110/hr
The role in one line
You create and evaluate research-level physics problems that test AI reasoning in your specific atomic, molecular, or optical subfield.
Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.
What you would do
- Design or solve research-level physics problems in your area that challenge genuine research thinking rather than formula recall
- Set up Hamiltonians, apply approximations, and justify boundary conditions using domain-specific methods from your publication record
- Write complete solutions with clear reasoning that another specialist can independently verify and critique
- Review AI-generated physics work and assess whether it demonstrates authentic research logic or superficial pattern-matching
- Document all work in LaTeX and Python with transparent mathematical justification
Who they are looking for
- PhD in atomic, molecular, optical physics or closely related field (hard requirement)
- Published papers demonstrating mastery of one specific phenomenon (3-5 recent papers with arXiv ID/DOI)
- First-author publication record strongly preferred; postdocs, research scientists, or junior faculty are ideal
- Comfortable using LaTeX, Python, SymPy, and Jupyter for your calculations and documentation
- English at B2+ level with strong written reasoning and mathematical communication skills
Skills this role asks for
What the interview is likely to probe
1.Phenomenon-specific problem construction
The benchmark tests whether you can pose questions that distinguish real research thinking from textbook solving; generic physics knowledge fails here.
Expect something like: “Design a research-level problem on your specific phenomenon where a model could easily fake understanding by pattern-matching textbook methods but would fail on a key physical insight.”
2.Justifying approximations and boundary conditions
Research physics lives in the reasoning behind simplifications; you need to defend why a dipole approximation or tight-binding assumption holds in your specific system.
Expect something like: “Walk me through a regime in your research where the standard approximation breaks down and how you handle that in a problem designed to test an AI's grasp of the limit.”
3.Publication-backed technical depth
Genuine expertise on the exact phenomenon you claim is essential; mastery of methods alone won't suffice without demonstrating publication in that specific area.
Expect something like: “Take one of your papers and describe a research problem you could derive from it that would expose whether an AI understands the method or just knows the name.”
4.Detecting physics reasoning versus mimicry
AI can sound plausible on physics; spotting where it skips the actual reasoning separates someone who reads AI work carefully from someone who assumes coherence.
Expect something like: “Show me an AI solution to a research problem that looks superficially correct but misses a physical constraint central to your phenomenon. How would you grade that?”
5.Documentation rigor for reproducibility
Your work trains AI; vague or incomplete writeups mislead the training dataset; research-grade requires another expert to verify independently.
Expect something like: “You found an error in an AI's approach to cavity QED that involves Lindblad evolution. Write the feedback you'd give explaining what the error is and why it matters physically.”
Exercise you may get
Given your research area and one of your published methods, design a research-level problem that would expose whether an AI understands the physics or merely patterns over similar solutions.
How to prepare
- Gather your 3-5 strongest first-author papers and map which specific methods and phenomena each demonstrates for quick reference
- Identify one research problem from your field where you can articulate the key physical insight a model must grasp to reason correctly
- Review the CritPt benchmark paper (arXiv:2509.26574) to understand how research-level problems differ from textbook ones
- Prepare one example where you caught a subtle error in published work or peer review that highlights the kind of reasoning depth the benchmark tests
Facts
- Pay
- $80–110/hr
- Commitment
- hourly
- Hours
- 10 per week
- Work arrangement
- remote · Remote
- Domain
- Life, Physical, and Social Science
- Posted
- 9/25/2026
- Open slots
- 3