Training Turk

Data Scientist Talent Network

$100–150/hr · Mercor · Hourly, 40 hours a week

You evaluate how well AI systems handle real data science problems, scoring their work against rigorous criteria.

What you would do

  • Design grading criteria specific to different data science tasks (e.g., exploratory analysis, modeling, experimentation)
  • Evaluate AI-generated or peer work against established standards with objective judgment
  • Write detailed rationales explaining scores and identifying technical errors or reasoning gaps
  • Independently verify reported findings by writing and running SQL queries against source data
  • Calibrate your scoring with senior reviewers to ensure consistency and defensibility

Who they want

  • 1+ years of hands-on data science work in a professional setting
  • Strong command of Python, SQL, and statistical modeling techniques
  • Deep understanding of machine learning workflows, experimentation design, and causal inference
  • Exceptional written communication: ability to explain technical findings clearly to non-specialists
  • Experience at a top technology company, research lab, quantitative fund, or equivalent organization

Main skills

Python and SQL proficiencyStatistical modeling fundamentalsMachine learning pipeline evaluation

What the interview asks about

  1. 1.Detecting hidden analytical errors

    A model might predict correctly by accident; identifying whether logic is sound matters more than whether output looks right.

    For example: “An analysis reports revenue down 12% in Q3, with SQL that left-joins transaction data to a calendar table. You notice the join key has nulls. What's your hypothesis and how do you verify it?”

  2. 2.Evaluating statistical rigor

    AI can apply techniques mechanically without understanding assumptions; you must catch cases where method doesn't match problem.

    For example: “An A/B test reports a p-value of 0.032, but sample sizes differ by 40% between variants. How do you assess whether the conclusion is defensible?”

  3. 3.Writing scoring rationales

    Scores only matter if someone downstream understands why; vague feedback blocks improvement and fails quality review.

    For example: “You scored a modeling notebook 6/10 but initially wrote 'poor feature engineering'. What concrete evidence and clearer language would your senior reviewer expect?”

  4. 4.Collaborating with calibration teams

    Your judgment is only trustworthy if it aligns with organizational standards; you must adjust and defend your reasoning.

    For example: “A senior reviewer scores a messy but correct analysis 8/10 while you scored 6/10 for documentation gaps. How do you discuss this disagreement?”

  5. 5.Independence in verification

    Trusting reported outputs without verification defeats evaluation; rebuilding analyses from scratch catches errors and confirms understanding.

    For example: “A notebook claims that 85% of customers return within 30 days. You reconstruct the query and get 78%. What details do you check before escalating?”

A task you may get

Review a synthetic machine learning pipeline on a classification task: score it against rubric criteria (correctness, code clarity, assumptions stated), write a 150-word justification, and write SQL to independently verify one reported metric.

How to prepare

  • Recall real data science projects where subtle errors created plausible wrong answers (e.g., off-by-one in time windows, misaligned definitions)
  • Review sample scoring rubrics used in academic or enterprise peer review to understand how technical depth translates to scores
  • Practice writing technical feedback: take a flawed analysis and document specific errors with reasoning, not just grades

The facts

Pay
$100–150/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Posted
7/29/2026
Places left
10

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.