$70–120/hr · Mercor · Part time, 40 hours a week
A machine learning practitioner who evaluates and judges the quality of AI model outputs for research teams.
What you would do
- Write realistic modeling problems based on production systems you have deployed, including data, objectives, and reference answers
- Grade model-generated code and evaluation logic to verify whether implementation matches stated goals and avoids methodological errors
- Identify approaches that train cleanly but would fail in production due to distribution shift or misaligned incentives
- Review ablation studies and experimental designs for flawed conclusions or hidden assumptions
Who they want
- Professional experience training or deploying machine learning systems in production environments
- Ability to communicate technical reasoning clearly in writing for asynchronous review
- Comfort with incomplete specifications and willingness to surface ambiguities before starting work
- Track record of troubleshooting models after they encounter real data or users
- Familiarity with evaluation pitfalls like metric gaming, label noise, and train-test mismatch
Main skills
What the interview asks about
1.Production failure identification
Distinguishing between models that look good on test data and those that will actually work at scale is core to this role and directly impacts research quality.
For example: “A model's time-series forecast has competitive validation loss but exploits seasonal noise patterns. How would you identify and explain this problem?”
2.Evaluation methodology
Many models generate plausible-sounding evaluation setups that leak information or reward the wrong behavior, and catching this requires both skepticism and clear reasoning about data flow.
For example: “A model proposes using stratified cross-validation to evaluate a churn prediction model built on weekly customer snapshots. The code looks clean. What methodological risk would you flag here and why?”
3.Objective alignment
Models may implement something related to the stated goal but not exactly what was requested, and spotting this misalignment is essential for evaluating code you did not write.
For example: “A client asks you to evaluate whether a model's code correctly implements a revenue-maximizing objective. The loss function minimizes error on high-revenue customers. What question would you ask before approving this as correct?”
4.Real system complexity
Experience with actual deployed models teaches you failure modes that pure theory does not, and this context is what research teams lack and are hiring for.
For example: “Describe a machine learning system you built or maintained where the offline evaluation metrics looked strong but the model's live performance surprised you. What gap did you discover and how would you write evaluation guidance to catch similar problems?”
5.Ambiguity resolution
Vague or conflicting task specifications are common, and how you clarify them affects the quality and scope of the work you deliver and how useful your feedback becomes.
For example: “A research team says they want you to grade a model's 'decision quality' on a complex task but provide no explicit success metric. How would you clarify the assignment before writing your evaluation?”
6.Constraint communication
Production models operate under real constraints like latency, memory, or computational budget that training problems often ignore, and flagging this mismatch helps research teams build usable models.
For example: “You write a modeling problem based on a customer support system where your solution trains in 2 hours. A model's approach trains in 20 minutes but requires retraining every 6 hours to stay accurate. How would you present this trade-off in your write-up?”
A task you may get
Write a modeling problem from production ML you've deployed, with objective, constraints, and success criteria. Grade a model approach against this rubric.
How to prepare
- Document a production ML system you deployed, noting the offline metrics that looked good and any surprises when the model went live
- Study three common evaluation pitfalls such as data leakage, label noise exploitation, and metric misalignment by reviewing papers or case studies
- Practice explaining a technical modeling failure to someone without ML experience, focusing on why it happened and what an evaluator should check
- Review your own past code or proposals and identify one decision that seemed sound during development but created production risk
The facts
- Pay
- $70–120/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.