$40–70/hr · Mercor · Part time, 40 hours a week
Evaluate AI reasoning quality by writing prompts, grading model outputs, and flagging where models fail to follow instructions or reason correctly.
What you would do
- Write prompts and reference answers that establish what good reasoning looks like on general knowledge questions
- Grade AI model outputs on whether answers are correct, complete, and actually follow the instructions given
- Identify logical contradictions, confidently wrong statements, and unsupported claims in AI-generated text
- Explain in writing precisely why a model response fails to meet standards or misses critical reasoning steps
- Flag ambiguous instructions and request clarification when prompts don't clearly define expected responses
Who they want
- Professional or academic background showing careful judgment and analytical thinking from any field
- Strong written communication skills, since most work involves explaining your reasoning clearly
- Willingness to engage with ill-defined problems and ability to identify and communicate ambiguous specifications constructively
- Ability to read precisely and identify gaps between what was asked and what was delivered
- No specific credential required - strong judgment and clear thinking matter more than specialized expertise
Main skills
What the interview asks about
1.Instruction comprehension and adherence checking
AI models often provide technically correct but irrelevant answers, so evaluators must catch when models answer a different question than what was asked.
For example: “A prompt asks for the best evidence supporting a historical claim. An AI provides a technically accurate description of the era but never mentions supporting evidence. How would you grade this response?”
2.Reasoning chain and logical consistency
Identifying whether conclusions actually follow from premises requires careful analysis of each logical step in an AI argument.
For example: “An AI argues that since Company A is larger than Company B, Company A's products are higher quality. Identify the logical flaw and explain why this reasoning fails.”
3.Completeness and sufficiency assessment
Evaluating whether AI responses fully address prompt requirements means identifying what's missing or incomplete despite being partially correct.
For example: “A prompt asks for three reasons why something matters and the implications of each. An AI provides three reasons but no implications. How would you document this incompleteness for a technical team?”
4.Contradictions and confidence errors
AI systems sometimes make contradictory statements within a single response, showing lack of internal consistency that must be caught and explained.
For example: “In a single response, an AI claims economic policy A reduces inflation but later says the same policy increases inflation. How would you flag and document this contradiction?”
A task you may get
Write a prompt on a general reasoning question with a reference answer showing good logic, then grade three sample AI responses identifying reasoning errors and gaps.
How to prepare
- Collect and analyze examples of AI model outputs on reasoning tasks to understand common error patterns and inconsistencies
- Develop templates for explaining AI reasoning failures clearly - what makes an explanation of an error most helpful to a technical team
- Review critical thinking frameworks and logical reasoning standards to calibrate your own judgment criteria
- Practice writing prompts that clearly define what constitutes a good answer and what context reasoning needs to account for
The facts
- Pay
- $40–70/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Miscellaneous
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.