$60–100/hr · Mercor · Full time, 40 hours a week
Senior business operations expert who reviews and trains AI systems to reason about real operations challenges through quality assessment, benchmark design, and domain standards.
What you would do
- Review AI-generated operations work and task outputs, identifying missing reasoning, thin assumptions, numerical inconsistencies, and answers that lack operating depth
- Write specification documents and golden solution sets that model how correct operations decisions are made in realistic, constraint-filled scenarios
- Design and build operations benchmarks and evaluation frameworks that measure whether AI systems genuinely understand operational tradeoffs
- Collaborate with research teams to calibrate standards and translate your operating judgment into explicit, teachable criteria for AI training
- Provide detailed feedback on model performance and help researchers understand where AI reasoning diverges from actual operating practice
Who they want
- Minimum four years managing a business, sales, or revenue operations function at a senior level, with real budget or team responsibility
- Deep specialization in at least one area: sales operations, revenue forecasting, quota design, CRM systems, pricing, supply chain, or process improvement
- Bachelor's degree required; MBA or equivalent advanced degree strongly preferred
- Hands-on experience using large language models in your work and the judgment to distinguish sound reasoning from plausible-sounding errors
- Must live in the Bay Area and commit 40 hours per week on-site for at least 6 months
Main skills
What the interview asks about
1.Spot operational reasoning flaws
AI often produces outputs that read convincingly but miss critical operating constraints, numbers that do not reconcile, or assumptions that do not hold in practice. Evaluators must catch these gaps before they corrupt training data.
For example: “An AI generates a sales territory plan that doubles headcount and predicts 40 percent revenue growth in 18 months. What operating details would you scrutinize to assess whether this plan is realistic, and what would signal the AI missed important constraints?”
2.Define correct operations solutions
Golden solutions must model how a seasoned operator actually makes tradeoff decisions under real constraints. AI learns from these examples, so they must be thorough and operationally sound.
For example: “Write a golden solution for a scenario where a business must cut costs 15 percent while maintaining revenue and satisfaction. What elements help AI understand competing priorities and operator tradeoffs?”
3.Design evaluation tasks
Good benchmarks isolate specific operations skills and force AI to apply reasoning, not memorize patterns. Poor benchmarks let AI pass without genuine understanding.
For example: “Design a benchmark task that tests whether AI can recognize when a quota assignment violates compensation policy, market capacity, or realistic attainment. What variations or constraints would you build in to prevent AI from gaming the scenario?”
4.Translate operating judgment
AI researchers need explicit criteria to calibrate training. Your job is converting intuitive operating decisions into clear, measurable standards that others can apply consistently.
For example: “You've spent years making pricing decisions. What explicit framework or checklist would you create to help a non-operator evaluate whether an AI's pricing recommendation accounts for margin, competitive positioning, and customer segments?”
5.Recognize AI gaps in constraint reasoning
AI systems often fail at operations because they miss implicit constraints or optimize for one metric while ignoring others. Experienced operators spot these patterns instantly.
For example: “An AI proposes centralizing inventory at one hub to cut costs 12 percent but does not address how this affects delivery times, supplier lead times, or regional demand swings. How would you provide feedback that helps the research team understand this AI gap?”
6.Assess numerical consistency
Operations decisions depend on numbers adding up and logic flowing. AI may miss when headcount, budget, revenue, or resource assumptions are internally inconsistent.
For example: “You review an AI recommendation to hire 50 salespeople, increase territory size by 30 percent, and expect 25 percent revenue growth. What numerical audit would you conduct, and what inconsistencies would you flag?”
A task you may get
Review an AI-generated operations task or solution in your domain (sales operations, revenue forecasting, supply chain, or process improvement). Identify quality gaps, score the reasoning against real operating standards, and draft specification improvements.
How to prepare
- Select one operations domain you know deeply. Map out 3-4 major decision types (e.g., quota setting, territory design). Note what AI must understand to reason well.
- Gather one realistic operations scenario you solved. Sketch a golden solution: what constraints did you consider, and how did you decide?
- Think of an AI output that sounded plausible but violated operating reality. What was wrong and how would you teach AI to recognize it?
- Practice explaining a nuanced operations decision you made to someone outside your field. This is the skill of translating tacit judgment into teachable criteria.
The facts
- Pay
- $60–100/hr
- Hours
- Full time, 40 hours a week
- Where
- Hybrid · Bay Area, CA
- Open to
- USA
- Field
- Business Operations
- Posted
- 8/19/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.