Training Turk

Expert Senior SWE

$150–210/hr · Mercor · Hourly

You evaluate and improve how frontier AI models perform on complex software engineering tasks, working closely with AI lab researchers.

What you would do

  • Assess language model outputs on coding, architecture, and system design tasks
  • Write evaluation test cases and code samples that surface model capability gaps
  • Review model reasoning on complex engineering problems and design tradeoffs
  • Provide detailed feedback on failures, including technical root causes
  • Contribute to prioritizing which SWE capabilities the model should improve next

Who they want

  • 10+ years of software engineering at top-tier US or European technology companies
  • Passionate about advancing AI model capabilities in complex technical domains
  • Able to commit approximately 20 hours per week for the engagement period
  • Experience with system design, distributed systems, or architectural thinking
  • Strong track record shipping production systems with high technical standards

Main skills

Large language model evaluationAI model capability assessmentSystem design and architecture

What the interview asks about

  1. 1.Model reasoning on architecture

    Language models often sound confident on system design but make subtle tradeoffs incorrectly; your judgment separates plausible-sounding responses from sound architecture.

    For example: “A model suggests an eventually-consistent cache with TTL for a high-frequency service. It sounds plausible but misses a key detail about your use case. What detail matters?”

  2. 2.Code quality judgment

    Models generate syntactically correct code that violates production standards (concurrency bugs, resource leaks, poor testability); you identify non-obvious quality failures.

    For example: “A model writes working Go code for a database connection pool that compiles and runs but has a subtle race condition under load. What pattern would you look for, and how would you test whether the model understands concurrency primitives?”

  3. 3.Problem decomposition insight

    Senior engineers decompose hard problems well; evaluating whether a model can break a complex task into solvable subproblems reveals reasoning depth.

    For example: “Asked to design a rate limiter serving millions of requests per second, a model's response is architecturally sound but omits how to handle clock skew across servers. Do you flag this as a gap or a reasonable scope boundary? Why?”

  4. 4.Failure mode analysis

    You need to trace why a model failed so researchers understand what training data or technique would help; generic 'this is wrong' feedback doesn't drive improvement.

    For example: “A model's parser code has correct pseudocode but fails on operator precedence. Is this memorization weakness, reasoning gap, or code generation issue? How do you distinguish?”

  5. 5.Research collaboration communication

    AI lab researchers need clear, precise feedback that grounds their data curation or training decisions; vague critiques waste their time.

    For example: “You've identified that a model struggles with lock-free data structures. How do you communicate this gap to researchers in a way that suggests what kind of training data or code examples might help improve performance?”

A task you may get

Evaluate three language model code outputs on progressively harder tasks: implementation task, system design problem, and architectural tradeoff question, with detailed analysis of why each succeeds or fails.

How to prepare

  • Prepare two or three specific software engineering challenges you've solved that required non-obvious decisions; they'll anchor your evaluation standards
  • Review recent advances in AI model capabilities for code and identify what still seems impossible; this frames where evaluation is most valuable
  • Prepare examples of clean versus poor production code to calibrate what 'high quality' means in your context

The facts

Pay
$150–210/hr
Hours
Hourly
Where
Remote
Field
Software Engineering
Project name
Dorado
Posted
7/19/2025
Places left
28

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.