Training Turk

SWE-Bench Task Auditor

$70–90/hr · Mercor · Hourly, 40 hours a week

A software engineer who audits AI model training benchmark tasks for correctness, reproducibility, and grading integrity.

What you would do

  • Review reference patches and implementation details for logical soundness and environmental correctness
  • Examine test harnesses, Docker configurations, and grading systems for reproducibility and fairness
  • Identify answer leakage, reward hacking, and unintended shortcuts that could inflate model scores
  • Assess repository-level tasks across multiple programming languages and ecosystems
  • Provide detailed, rubric-based feedback on task quality and suggest fixes

Who they want

  • 3+ years of professional software engineering with real open-source contributions or maintainer roles
  • Strong skills in code review, patch analysis, and detecting subtle implementation issues
  • Strong Python expertise combined with proficiency in at least one additional language such as Java, Go, TypeScript, or C++
  • Experience auditing test infrastructure, Docker configurations, and automated grading systems
  • Familiarity with SWE-Bench or similar repository-level benchmarking frameworks

Main skills

Code review and auditingDocker and test isolationReference patch evaluation

What the interview asks about

  1. 1.Reference patch evaluation

    A flawed reference patch directly corrupts benchmark results, causing valid solutions to fail or invalid ones to pass. Spotting these issues early prevents months of wasted evaluation.

    For example: “A Python package reference patch modifies 15 files. Model solutions pass at 60%, reference at 100%. What patch elements would you examine first?”

  2. 2.Answer leakage detection

    Models can exploit unintended information in test harnesses - timing signals, error messages, or harness structure - to fake correct answers without genuine problem-solving. This invalidates benchmark results.

    For example: “A Go concurrency benchmark compares model output against reference code. How would you verify the evaluation system doesn't accidentally leak solution structure?”

  3. 3.Test infrastructure isolation

    Reproduction failures and environmental dependencies account for most invalid benchmark tasks. Verifying Docker isolation and cross-environment consistency ensures fair comparison of model performance.

    For example: “A TypeScript benchmark shows references at 0.5s but submissions at 8s with identical output. How would you distinguish infrastructure issues from correctness problems?”

  4. 4.Multi-ecosystem grading consistency

    Benchmarks assess models across Python, Java, Go, TypeScript, and C++. Language-specific issues in test harnesses or reference patches can unfairly penalize some ecosystems and inflate performance in others.

    For example: “A C++ task reference solution uses a header outside the repo root. Is this legitimate testing or a grading flaw? How would you determine this?”

  5. 5.Reproducibility verification

    Benchmarks must produce consistent results across environments and dependency versions. Hidden reproducibility issues undermine the validity of model rankings and make debugging impossible.

    For example: “A Django package benchmark patch regenerates database migrations. How would you verify reproducibility across Python versions, dependencies, and fresh database states?”

A task you may get

Audit a benchmark task: examine the reference patch, test harness, Docker configuration, and grading system. Identify correctness issues and answer leakage.

How to prepare

  • Practice reviewing code patches for subtle environment dependencies, version assumptions, and unintended side effects
  • Study Docker best practices for isolated testing and understand how test infrastructure can accidentally leak information
  • Explore SWE-Bench or HumanEval repository structures to understand benchmark task design patterns
  • Review examples of answer leakage and reward hacking in machine learning evaluation systems

The facts

Pay
$70–90/hr
Hours
Hourly, 40 hours a week
Where
Remote · Remote — United States
Open to
USA
Field
Software Engineering
Posted
8/28/2026
Places left
3

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.