$100–150/hr · Mercor · Hourly, 40 hours a week
You evaluate how well AI systems handle real data science problems, scoring their work against rigorous criteria.
What you would do
- Design grading criteria specific to different data science tasks (e.g., exploratory analysis, modeling, experimentation)
- Evaluate AI-generated or peer work against established standards with objective judgment
- Write detailed rationales explaining scores and identifying technical errors or reasoning gaps
- Independently verify reported findings by writing and running SQL queries against source data
- Calibrate your scoring with senior reviewers to ensure consistency and defensibility
Who they want
- 1+ years of hands-on data science work in a professional setting
- Strong command of Python, SQL, and statistical modeling techniques
- Deep understanding of machine learning workflows, experimentation design, and causal inference
- Exceptional written communication: ability to explain technical findings clearly to non-specialists
- Experience at a top technology company, research lab, quantitative fund, or equivalent organization
Main skills
What the interview asks about
1.Detecting hidden analytical errors
A model might predict correctly by accident; identifying whether logic is sound matters more than whether output looks right.
For example: “An analysis reports revenue down 12% in Q3, with SQL that left-joins transaction data to a calendar table. You notice the join key has nulls. What's your hypothesis and how do you verify it?”
2.Evaluating statistical rigor
AI can apply techniques mechanically without understanding assumptions; you must catch cases where method doesn't match problem.
For example: “An A/B test reports a p-value of 0.032, but sample sizes differ by 40% between variants. How do you assess whether the conclusion is defensible?”
3.Writing scoring rationales
Scores only matter if someone downstream understands why; vague feedback blocks improvement and fails quality review.
For example: “You scored a modeling notebook 6/10 but initially wrote 'poor feature engineering'. What concrete evidence and clearer language would your senior reviewer expect?”
4.Collaborating with calibration teams
Your judgment is only trustworthy if it aligns with organizational standards; you must adjust and defend your reasoning.
For example: “A senior reviewer scores a messy but correct analysis 8/10 while you scored 6/10 for documentation gaps. How do you discuss this disagreement?”
5.Independence in verification
Trusting reported outputs without verification defeats evaluation; rebuilding analyses from scratch catches errors and confirms understanding.
For example: “A notebook claims that 85% of customers return within 30 days. You reconstruct the query and get 78%. What details do you check before escalating?”
A task you may get
Review a synthetic machine learning pipeline on a classification task: score it against rubric criteria (correctness, code clarity, assumptions stated), write a 150-word justification, and write SQL to independently verify one reported metric.
How to prepare
- Recall real data science projects where subtle errors created plausible wrong answers (e.g., off-by-one in time windows, misaligned definitions)
- Review sample scoring rubrics used in academic or enterprise peer review to understand how technical depth translates to scores
- Practice writing technical feedback: take a flawed analysis and document specific errors with reasoning, not just grades
The facts
- Pay
- $100–150/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Posted
- 7/29/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.