$60–70/hr · Mercor · Hourly, 40 hours a week
A safety specialist who evaluates AI model outputs for compliance, safety, and alignment with organizational policies.
What you would do
- Evaluate AI-generated responses for safety compliance, factual accuracy, and adherence to organizational policies
- Review content involving high-risk domains including misinformation, persuasion tactics, violence, and biosecurity threats
- Identify unsafe outputs, hallucinated information, reasoning failures, and policy violations with detailed explanations
- Apply and refine evaluation rubrics used for reinforcement learning and safety benchmarking processes
Who they want
- Bachelor's degree in journalism, communications, psychology, public policy, law, computer science, or related discipline
- 5+ years professional experience in AI safety, trust and safety, journalism, public policy, security, or related field
- Excellent written communication, critical thinking, and analytical reasoning skills
- Capacity to regularly assess complex and sensitive situations with discernment
- Preferred: direct AI safety, RLHF, content moderation, or evaluation rubric development experience
Main skills
What the interview asks about
1.Policy violation identification
Accurately flagging policy violations requires understanding organizational safety rules and recognizing subtle violations that automated systems might miss.
For example: “An AI response appears helpful on the surface but uses subtle persuasion tactics to influence political opinion without disclosing intent. How would you classify this against a policy that permits factual discussion but prohibits covert persuasion?”
2.Hallucination detection accuracy
Distinguishing genuine hallucinations from rare but plausible outputs is critical; false positives undermine training, while missed hallucinations create safety gaps.
For example: “An AI response cites a scientific study with realistic formatting and details, but you suspect it's fabricated. You cannot look it up in the evaluation window. What approach would you take to assess whether this is a hallucination?”
3.Gray-area judgment calls
Many safety scenarios lack clear-cut answers; evaluators must apply judgment consistently while acknowledging genuine ambiguity in policy-sensitive topics.
For example: “A response contains technically accurate information about a cybersecurity vulnerability but could potentially enable malicious use. The policy explicitly permits educational security content. How would you evaluate this?”
4.Reasoning failure analysis
Identifying when models reach correct conclusions through flawed logic or incorrect intermediate steps helps teams improve model reasoning quality, not just final answers.
For example: “An AI response arrives at the correct policy recommendation but uses several questionable logical leaps and mischaracterizes a prior precedent in the reasoning. How would you structure feedback?”
5.Rubric consistency application
Consistent rubric application ensures training data quality and prevents bias; inconsistency can introduce systematic errors that models learn from.
For example: “You've evaluated 50 responses using a safety rubric. You notice your severity ratings for ambiguous cases drifted from strict to lenient. How would you recalibrate consistency?”
A task you may get
Evaluate a set of AI responses across varied risk domains, identifying policy violations, hallucinations, and reasoning failures, then providing structured feedback.
How to prepare
- Review current AI safety policies, common violation patterns, and evaluation rubric frameworks
- Study types of AI hallucinations and reasoning failures to develop recognition skills
- Research policy-sensitive topics like misinformation, persuasion, biosecurity to build domain context
- Practice giving structured, actionable feedback that identifies problems and explains significance
The facts
- Pay
- $60–70/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Open to
- USA, DNK, EST, FIN, ISL, IRL, LVA, LTU, NOR, SWE, AUT, BEL, FRA, DEU, LIE, LUX, MCO, NLD, CHE, GBR, ALB, BIH, HRV, GRC, ITA, XKX, MLT, MKD, PRT, SMR, SRB, SVN, ESP, BGR, CZE, HUN, MDA, POL, ROU, SVK, USA, DNK, EST, FIN, ISL, IRL, LVA, LTU, NOR, SWE, AUT, BEL, FRA, DEU, LIE, LUX, MCO, NLD, CHE, GBR, ALB, BIH, HRV, GRC, ITA, XKX, MLT, MKD, PRT, SMR, SRB, SVN, ESP, BGR, CZE, HUN, MDA, POL, ROU, SVK
- Field
- Life, Physical, and Social Science
- Project name
- AIUC
- Posted
- 7/16/2026
- Places left
- 4
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.