$2k/task
The role in one line
A technical expert designs a challenging, realistic task that tests AI agents on expert-level work in a scientific domain.
Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.
What you would do
- Propose a difficult, open-ended problem grounded in authentic work within your specialization
- Write a clear task description that sets up the challenge for AI agents
- Design an automated scoring rubric with partial credit for incomplete or partial success
- Document the rationale for why this task effectively measures expert capability
Who they are looking for
- STEM expert with depth sufficient to identify and articulate genuinely challenging problems in your field
- Experience recognizing what makes a problem difficult versus routine or trivial
- Ability to think about evaluation criteria and how to score quality of solutions
- Comfort with tight timelines and producing polished, deployment-ready work within days
Skills this role asks for
What the interview is likely to probe
1.Problem identification and framing
The foundation of a good task is correctly identifying what constitutes expert-level difficulty; weak problem selection ruins the whole task.
Expect something like: “Describe one challenging problem in your field that doesn't have an obvious solution. Why would this frustrate a moderately skilled AI agent but be tractable to an expert?”
2.Scoring and evaluation logic
Open-ended tasks need fair, unambiguous scoring; poor scoring criteria make the task either too easy or unchallengeable.
Expect something like: “If you were scoring AI agent attempts at your task on a scale from 0 to 1, what would a 0.3 solution look like versus a 0.7 solution? Walk me through your scoring logic.”
3.Task clarity and specification
An AI agent must understand the task without ambiguity; vague instructions produce confused attempts and unmeasurable results.
Expect something like: “Write the opening paragraph an AI agent would read to understand your task. What specific constraints or requirements would you include, and which are intentionally left open?”
4.Complexity versus feasibility
Tasks must challenge AI agents without becoming computationally impossible; balancing this requires deep technical judgment.
Expect something like: “How would you adjust your task if early AI attempts succeeded trivially? What aspects would you make harder, and what would you preserve?”
5.Real-world grounding
Tasks that exist only in theory don't measure genuine expert capability; authentic problems ensure the evaluation is meaningful.
Expect something like: “Describe how your task mirrors actual work a professional in your field does. What real-world constraints or contexts motivated this design?”
Exercise you may get
Outline one open-ended technical task in your field: describe the challenge in 2-3 sentences, specify what a partial-credit scoring rubric would assess, and explain why this task would reveal capability gaps in AI agents.
How to prepare
- Identify 3-4 difficult problems you've encountered or observed in your professional work that lack formulaic solutions
- Think about how scoring frameworks distinguish mediocre from excellent attempts; prepare examples of work at different quality tiers
- Consider what makes a problem genuinely open-ended versus having a hidden single correct answer
- Reflect on what you'd want an AI system to understand or demonstrate when solving a complex problem in your specialty
Facts
- Pay
- $2k/task
- Commitment
- task-based
- Work arrangement
- remote · Remote
- Posted
- 10/2/2026