$80–110/hr
The role in one line
PhD physicist with published research on specific high-energy theory phenomena creating and evaluating research-grade benchmark problems for AI model assessment.
Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.
What you would do
- Create research-level physics problems testing genuine reasoning in your published subfield, with complete problem statement, assumptions, and reference solutions
- Solve research-grade challenges in your area of expertise, documenting your reasoning for independent verification
- Review and audit completed physics work from peers to confirm correctness, rigor, and valid reasoning chains
- Communicate all findings in writing clear enough for qualified reviewers in your subfield to understand and verify independently
- Match with project assignments aligned to your publication record and specific research methods
Who they are looking for
- Doctoral degree in theoretical physics or a closely related field with active research experience
- Published work on specific phenomena (not adjacent topics) proven via arXiv ID or DOI in your papers
- Postdoctoral researcher, research scientist, junior faculty, or senior PhD student with strong first-author record
- Verifiable publication record with 3-5 representative papers from the last 5 years; first author preferred
- B2-or-above English proficiency for clear written reasoning and technical communication
What the interview is likely to probe
1.Phenomenon-specific mastery
The benchmark requires deep knowledge of particular physics you've published on; shallow familiarity with the broader field leads to weak problems and unreliable reviews.
Expect something like: “You've published on AdS/BCFT correspondence. Describe a research-level problem that tests understanding of end-of-the-world branes specifically, not just AdS/CFT broadly.”
2.Problem design for genuine reasoning
Good benchmarks force models to think, not pattern-match; your problem must expose the gap between memorization and real understanding.
Expect something like: “You're designing a problem in your subfield. What feature would you include that a model trained on textbooks might miss, but any specialist would know to check?”
3.Documenting reasoning for peer review
Your write-up is how another specialist verifies your work; unclear logic means they cannot trust your solution or learn your approach.
Expect something like: “You solved a research problem in your area and need to document it so a peer can check your work independently. What level of detail would you include, and why?”
4.Identifying reasoning errors under assumptions
Research physics is assumption-dependent; spotting when a solution breaks under regime changes or assumption violation shows deep understanding.
Expect something like: “In your field, you're reviewing another physicist's solution. It reaches the correct numerical answer but relies on an approximation valid only in the weak-coupling limit. The problem specifies strong coupling. How would you report this in your review?”
Exercise you may get
Author a research-level physics problem in your published subfield with complete problem statement, assumptions, regime, methods, and reference solution.
How to prepare
- Gather your 3-5 most representative papers with arXiv IDs or DOIs ready to verify authorship and method demonstrating
- Review the CritPt benchmark paper (arXiv:2509.26574) to understand the quality and depth expected
- Identify 2-3 specific physics phenomena in your area where frontier AI might struggle and prepare to discuss what makes them hard
Facts
- Pay
- $80–110/hr
- Commitment
- hourly
- Hours
- 10 per week
- Work arrangement
- remote · Remote
- Domain
- Life, Physical, and Social Science
- Posted
- 9/25/2026
- Open slots
- 3