$70–150/hr · Mercor · Part time, 40 hours a week
You evaluate AI-generated code and architecture against production engineering standards and frontier AI training requirements.
What you would do
- Write test problems sourced from real systems you built, with full constraints and correct solutions
- Grade model code on edge case handling, concurrency safety, and cost-awareness
- Review system designs for scalability, testability, and operational feasibility
- Flag code that passes static analysis but fails under realistic load or failure scenarios
- Justify each critique in writing for training data and model evaluation
Who they want
- 4+ years shipping production software with hands-on responsibility for deployed systems
- Depth in at least one backend stack (Java, Go, Node.js) or full-stack (React, TypeScript, Next.js)
- Experience owning code through debugging failures caused by concurrency, scale, or infrastructure limitations
- Clear technical writing: you explain architectural tradeoffs to teams without code access
- Comfort with ambiguity and willingness to ask clarifying questions before assessing fuzzy specifications
Main skills
What the interview asks about
1.Spotting production failure modes
Many AI systems produce syntactically correct code that fails under real constraints; you must distinguish surface correctness from operational soundness.
For example: “A language model generated a caching layer using unbounded in-memory maps. Walk me through how you'd critique this for a service that needs to stay under 512 MB, run for days, and handle 1000 requests per second.”
2.Judging architectural overdesign
A model may generate solutions that are correct but unnecessarily complex; discerning developers know which simplifications are safe without sacrificing durability.
For example: “A model proposed a 4-layer abstraction for a single data fetch operation. How would you decide whether those layers add resilience and testability or just obscure the actual problem?”
3.Explaining design flaws clearly
Your role is to train AI on subtlety: you must articulate why something is wrong in terms a non-specialist reviewer can evaluate and learn from.
For example: “A model's test suite only validates happy paths. Write the critique you'd submit: why does this fail the standard, and what specific test case changes would fix it?”
4.Recognizing when you lack context
Engineering problems are often underspecified; pushing back on ambiguous requirements is part of producing usable training data.
For example: “You're asked to grade code for a system that 'needs to be fast,' with no throughput target or latency budget. How would you respond before evaluating the submission?”
A task you may get
Write a problem statement for one piece of production code you've owned: the requirements, key constraints, one serious failure mode that happened in production, and the final solution.
How to prepare
- Prepare 2-3 production incidents where you diagnosed failures caused by concurrency, caching, or scale
- Jot down 3 instances where a simple design change cut complexity without sacrificing correctness
- Collect examples of design patterns that seemed clever but created debugging nightmares
- Brush up on your stack's concurrency model and typical performance bottlenecks
The facts
- Pay
- $70–150/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Software Engineering
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.