$65–105/hr · Mercor · Full time, 40 hours a week
A senior software engineer vets AI model quality on engineering tasks and writes the specifications defining what correct engineering work looks like.
What you would do
- Review engineering tasks and model-generated solutions, identifying subtle errors, missed edge cases, and reasoning gaps that pass surface inspection but fail under scrutiny
- Write instruction specifications and golden solutions that encode your engineering judgment into explicit, measurable standards
- Design benchmark tasks that reflect real production constraints and test genuine improvement in model reasoning, not just pattern matching
Who they want
- Bachelor's or advanced degree in Computer Science or related engineering discipline; Master's or PhD preferred
- 4+ years of professional production engineering shipping real systems at reputable organizations; senior-level progression or equivalent impact
- Genuine specialization in at least one area: distributed systems, security, embedded systems, ML infrastructure, platform engineering, or quality engineering
- Evidence of engineering achievement: shipped products, granted patents, open source projects, or technical publications
- Comfort with LLMs and judgment to evaluate well-reasoned versus specious technical arguments
Main skills
What the interview asks about
1.Spotting subtle code defects
Models can generate code that runs but has subtle logic errors, unsafe assumptions, or performance traps; catching these requires hands-on engineering judgment.
For example: “An AI proposes a distributed cache invalidation approach that handles the happy path correctly but ignores network partitions. Your review needs to explain why this design fails and what a production system requires.”
2.Translating tacit knowledge into criteria
Benchmark quality depends on encoding your experience so others can apply it consistently and the model learns principles, not just your particular solutions.
For example: “You're writing the spec for a distributed system design task. Describe the non-obvious constraints (operational complexity, failure modes, monitoring burden) that textbook designs often overlook but your experience taught you are critical.”
3.Identifying when models confabulate
LLMs confidently generate plausible-sounding answers that are completely wrong; distinguishing hallucination from correct reasoning is the core of your role.
For example: “An AI explains a system architecture decision with coherent reasoning but the conclusion contradicts real constraints in the domain. How do you structure your feedback so researchers understand the error and improve the model?”
4.Designing task difficulty progression
Benchmarks need difficulty gradation to measure genuine improvement; too easy and models appear smarter than they are, too hard and nothing registers as progress.
For example: “You're building a benchmark on API design. Propose three tasks at increasing difficulty that test different aspects (naming consistency, backward compatibility, extension points) without requiring arbitrary domain knowledge.”
5.Evaluating engineering rigor
Production systems fail when engineers optimize for the wrong metrics - your job is defining what rigorous engineering means for AI training, not just correctness.
For example: “Two model responses both produce working code. One is minimal and elegant; the other is verbose but defensive with extensive error handling. Which is better for an AI training benchmark and why?”
A task you may get
Design a 3-task engineering benchmark in your specialization area, writing specifications, golden solutions, and clear grading rubrics that measure whether the model actually improved versus memorizing patterns.
How to prepare
- Review 2-3 recent technical blogs or papers from leading engineering teams showing real design decisions and trade-offs you'd expect an AI model to reason through
- Document one complex system you shipped, capturing the non-obvious constraints and failure modes you encountered that textbooks don't cover
- Prepare a short written example of engineering code review feedback that explains a subtle defect in a way that teaches principle, not just points out error
The facts
- Pay
- $65–105/hr
- Hours
- Full time, 40 hours a week
- Where
- Hybrid · Bay Area, CA
- Open to
- USA
- Field
- Software Engineering
- Posted
- 8/19/2026
- Places left
- 9
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.