Training Turk

Data Analysis Expert

$70–120/hr · Mercor · Part time, 40 hours a week

Data analyst who grades AI model outputs on analysis quality, spotting both technical errors and subtle biases that mislead decisions.

What you would do

  • Write analysis problems from work you actually shipped, including the data, question, constraints, and reference answer
  • Grade model-written analyses and SQL pipelines on correctness and whether they truly answer the question
  • Identify results that are technically valid but misleading due to selection bias, leakage, or wrong metrics
  • Assess join logic, fanout, and metric computation for accuracy and business sense
  • Provide detailed feedback on model strengths and failure modes

Who they want

  • Experience spanning data engineering, analytics, statistics, or machine learning
  • Proven track record of shipping analyses that drove actual business decisions
  • Strong written communication skills to explain reasoning clearly and thoroughly
  • Comfort handling ambiguous work; ability to flag instructions that lack clarity
  • Ability to spot when a result is technically correct but conceptually misleading

Main skills

Data analysis evaluationSQL query assessmentPipeline quality review

What the interview asks about

  1. 1.Query correctness and metric computation

    Technically correct SQL can compute the wrong thing; evaluating queries separates real analysts from pattern-matchers.

    For example: “A model writes a query to compute monthly churn rate by summing up churned customers and dividing by total customers. The query groups by month but includes customers who churned before the month started. Is the metric correct, and why?”

  2. 2.Bias and leakage spotting

    Subtle data errors lead to models that seem to work but fail in production; catching them in training prevents real-world harm.

    For example: “A model built a predictive model for next-quarter revenue using a feature that is only available after quarter-end. The model's test accuracy is 92%. What's the actual problem here?”

  3. 3.Reasoning clarity and communication

    A correct answer explained poorly is often worth less than a near-correct answer with clear logic.

    For example: “The model arrives at the right revenue forecast but its explanation conflates correlation with causation and doesn't mention the key seasonality driver. How do you rate the work?”

  4. 4.Join and aggregation logic

    Most real analytics mistakes come from fanout and mismatched keys; models struggle with this especially.

    For example: “You're asked to compute average order value per customer this quarter. The model joins orders to customers, then to line items, then sums revenue and divides by distinct customers. What's wrong with this approach?”

  5. 5.Metric versus outcome tracking

    Metrics are proxies; a model that optimizes the wrong metric is useless even if it's statistically sound.

    For example: “A model is asked to measure product quality. It computes a metric based on user ratings. But you know user ratings are biased toward power users and new users don't rate. How do you flag this as a problem?”

A task you may get

Write a short analysis problem using a real dataset you've worked with, including the business question, data description, and reference answer. Then review a sample model output and grade it.

How to prepare

  • Document a past analysis you shipped: the question, data, methodology, and outcome it drove
  • Reflect on one time a technically correct analysis turned out to be misleading in practice; understand why
  • Gather sample SQL or Python analytics code and review for correctness and reasoning clarity

The facts

Pay
$70–120/hr
Hours
Part time, 40 hours a week
Where
Remote
Field
Data Analysis
Role type
Talent network
Posted
9/15/2026

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.