$40–75/hr · Mercor · Part time, 40 hours a week
Evaluate advanced AI systems' mathematical capabilities by creating problems and assessing the rigor and correctness of generated proofs.
What you would do
- Compose original research or competition-difficulty problems with complete proofs and reference solutions
- Analyze model-generated proofs to verify logical validity at each step
- Identify reasoning gaps, unsupported invocations of theorems, and instances where edge cases are missed
- Distinguish between mathematically correct solutions and answers reached via invalid arguments
- Document detailed feedback on proof structure and reasoning quality
Who they want
- Graduate degree in mathematics or equivalent research/competition-level background
- Ability to write precisely and communicate complex mathematical reasoning clearly
- Tolerance for ambiguous or evolving project specifications and willingness to flag unclear instructions
- Experience judging mathematical rigor and distinguishing valid from invalid proofs
- Comfort working independently in remote settings on flexible schedules
Main skills
What the interview asks about
1.Verifying proof validity step-by-step
Your job is to catch subtle flaws; the interviewer needs to see that you can trace through a proof carefully and identify exactly where logic breaks.
For example: “A model's proof claims to establish convergence by applying the dominated convergence theorem but the conditions stated in the premises don't satisfy the theorem's hypotheses. How would you write feedback that explains this error precisely?”
2.Composing problems at research difficulty
The problems you create become the grading rubric for frontier models, so they must be genuinely challenging and unambiguously solvable.
For example: “Design a research-level problem in functional analysis that tests edge-case handling. Provide the complete solution and explain why a model might plausibly make an error on this problem.”
3.Distinguishing correct answer from correct proof
A model might guess the right answer or use an invalid shortcut; you must catch when the reasoning itself is flawed regardless of the final result.
For example: “A model correctly determines that a series converges but uses an informal appeal to 'intuition about growth rates' instead of applying a convergence test. How would you categorize and critique this?”
4.Responding to vague or evolving specifications
Pool-based work means instructions may be incomplete; the interviewer assesses whether you'll proceed carefully or get stuck.
For example: “You're asked to grade model-generated solutions to 'abstract algebra problems' but no difficulty level, topic, or success criteria are specified. What would you ask or assume?”
A task you may get
Compose one research-level mathematics problem in your specialty area, provide a complete proof, and identify one subtle mistake a model might plausibly make when solving it.
How to prepare
- Review recent publications in your mathematical field to ensure you're calibrated to frontier difficulty
- Practice writing precise mathematical exposition and identifying logical gaps in published proofs
- Collect examples of plausible but invalid mathematical arguments to sharpen your error-detection intuition
- Prepare to discuss a recent complex proof you've studied and any subtle errors or gaps you found in it
The facts
- Pay
- $40–75/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.