$20–80/hr
The role in one line
STEM specialist evaluating how AI systems reason about technical and scientific problems, ensuring accuracy and rigor in frontier domains.
Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.
What you would do
- Review AI responses to technical and scientific questions across physics, mathematics, chemistry, engineering, or data science domains
- Verify AI outputs against authoritative sources and first-principles reasoning, identifying errors and misconceptions
- Assess whether AI-generated data visualizations, calculations, and analyses are scientifically sound
- Provide structured feedback on AI reasoning gaps and areas needing improvement
- Contribute your expertise to help AI systems learn to handle complex technical problems correctly
Who they are looking for
- Advanced degree or substantial professional background in engineering, physics, mathematics, chemistry, software engineering, or data science
- Strong written English communication with clear, precise technical writing
- Critical judgment applied to verification - checking work rigorously rather than accepting plausible-sounding answers
- Data analysis and visualization experience preferred, showing ability to interpret quantitative information
- Availability for minimum 5 hours weekly on a flexible hourly schedule
Skills this role asks for
What the interview is likely to probe
1.Verifying technical correctness beyond surface plausibility
Plausible-sounding but incorrect answers are the most dangerous AI failures. Your role is catching subtle violations of fundamental principles that an untrained reviewer might miss.
Expect something like: “An AI system solves a physics problem and arrives at the correct numerical answer using an approach that violates energy conservation principles. The method happens to work for this case but is fundamentally wrong. How would you evaluate this response?”
2.Domain-specific error identification
General feedback isn't useful - AI researchers need to know specifically what principle was violated, what assumptions were incorrect, and why it matters in your field.
Expect something like: “An AI recommends a metal for high-temperature use. The choice is reasonable, but it ignores thermal fatigue that would fail it in weeks. How would you explain this gap?”
3.Data interpretation and visualization critique
AI systems generate charts, analyses, and summaries. You need to spot when visualizations obscure important patterns or when statistical conclusions don't follow from the data shown.
Expect something like: “An AI generates a chart showing experimental results with error bars so large they overlap completely. The AI then draws confident conclusions about statistically significant differences. What's the problem and how would you rate this work?”
4.Reasoning process versus answer correctness
Right answers for wrong reasons are still wrong in science. AI training improves when you can articulate whether an error is conceptual misunderstanding or computational mistake.
Expect something like: “An AI gets a chemistry problem right but justifies the answer using incorrect reaction mechanisms. The right product forms, but the reasoning path doesn't reflect what actually happens. How significant is this for model training?”
5.Distinguishing ambiguity from error
Not every AI gap is a failure - some questions are genuinely ambiguous or depend on assumptions. You need to recognize when the AI's interpretation is defensible versus clearly wrong.
Expect something like: “A mathematical optimization problem could have multiple correct solutions depending on constraints. An AI picks one and labels it optimal without discussing other possibilities. Is this an error, incomplete reasoning, or acceptable for the context?”
Exercise you may get
An AI analyzes a scientific dataset and draws conclusions. Review for technical accuracy, data interpretation, visualization quality, and reasoning. Identify three issues and explain their significance.
How to prepare
- Review recent papers or authoritative references in your specialty to refresh knowledge of current best practices and frontier concepts
- Prepare examples from your professional experience where a plausible-sounding approach failed because of subtle principle violations
- Practice articulating your domain expertise in clear, technical writing that researchers could use to improve AI systems
- Develop a personal checklist of common errors or misconceptions in your field that you'd watch for in AI outputs
Facts
- Pay
- $20–80/hr
- Commitment
- hourly
- Hours
- 40 per week
- Work arrangement
- remote · Remote
- Posted
- 9/18/2026