Training Turk

Mercor listing

High Energy Physicist (PhD)

$80–110/hr

The role in one line

PhD physicist with published research on specific high-energy theory phenomena creating and evaluating research-grade benchmark problems for AI model assessment.

Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.

What you would do

  • Create research-level physics problems testing genuine reasoning in your published subfield, with complete problem statement, assumptions, and reference solutions
  • Solve research-grade challenges in your area of expertise, documenting your reasoning for independent verification
  • Review and audit completed physics work from peers to confirm correctness, rigor, and valid reasoning chains
  • Communicate all findings in writing clear enough for qualified reviewers in your subfield to understand and verify independently
  • Match with project assignments aligned to your publication record and specific research methods

Who they are looking for

  • Doctoral degree in theoretical physics or a closely related field with active research experience
  • Published work on specific phenomena (not adjacent topics) proven via arXiv ID or DOI in your papers
  • Postdoctoral researcher, research scientist, junior faculty, or senior PhD student with strong first-author record
  • Verifiable publication record with 3-5 representative papers from the last 5 years; first author preferred
  • B2-or-above English proficiency for clear written reasoning and technical communication

What the interview is likely to probe

  1. 1.Phenomenon-specific mastery

    The benchmark requires deep knowledge of particular physics you've published on; shallow familiarity with the broader field leads to weak problems and unreliable reviews.

    Expect something like: “You've published on AdS/BCFT correspondence. Describe a research-level problem that tests understanding of end-of-the-world branes specifically, not just AdS/CFT broadly.”

  2. 2.Problem design for genuine reasoning

    Good benchmarks force models to think, not pattern-match; your problem must expose the gap between memorization and real understanding.

    Expect something like: “You're designing a problem in your subfield. What feature would you include that a model trained on textbooks might miss, but any specialist would know to check?”

  3. 3.Documenting reasoning for peer review

    Your write-up is how another specialist verifies your work; unclear logic means they cannot trust your solution or learn your approach.

    Expect something like: “You solved a research problem in your area and need to document it so a peer can check your work independently. What level of detail would you include, and why?”

  4. 4.Identifying reasoning errors under assumptions

    Research physics is assumption-dependent; spotting when a solution breaks under regime changes or assumption violation shows deep understanding.

    Expect something like: “In your field, you're reviewing another physicist's solution. It reaches the correct numerical answer but relies on an approximation valid only in the weak-coupling limit. The problem specifies strong coupling. How would you report this in your review?”

Exercise you may get

Author a research-level physics problem in your published subfield with complete problem statement, assumptions, regime, methods, and reference solution.

How to prepare

  • Gather your 3-5 most representative papers with arXiv IDs or DOIs ready to verify authorship and method demonstrating
  • Review the CritPt benchmark paper (arXiv:2509.26574) to understand the quality and depth expected
  • Identify 2-3 specific physics phenomena in your area where frontier AI might struggle and prepare to discuss what makes them hard

Facts

Pay
$80–110/hr
Commitment
hourly
Hours
10 per week
Work arrangement
remote · Remote
Domain
Life, Physical, and Social Science
Posted
9/25/2026
Open slots
3