Training Turk

AI Safety Practitioner

$60–70/hr · Mercor · Hourly, 40 hours a week

A safety specialist who evaluates AI model outputs for compliance, safety, and alignment with organizational policies.

What you would do

  • Evaluate AI-generated responses for safety compliance, factual accuracy, and adherence to organizational policies
  • Review content involving high-risk domains including misinformation, persuasion tactics, violence, and biosecurity threats
  • Identify unsafe outputs, hallucinated information, reasoning failures, and policy violations with detailed explanations
  • Apply and refine evaluation rubrics used for reinforcement learning and safety benchmarking processes

Who they want

  • Bachelor's degree in journalism, communications, psychology, public policy, law, computer science, or related discipline
  • 5+ years professional experience in AI safety, trust and safety, journalism, public policy, security, or related field
  • Excellent written communication, critical thinking, and analytical reasoning skills
  • Capacity to regularly assess complex and sensitive situations with discernment
  • Preferred: direct AI safety, RLHF, content moderation, or evaluation rubric development experience

Main skills

AI safety evaluationPolicy compliance assessmentRLHF and sft evaluation

What the interview asks about

  1. 1.Policy violation identification

    Accurately flagging policy violations requires understanding organizational safety rules and recognizing subtle violations that automated systems might miss.

    For example: “An AI response appears helpful on the surface but uses subtle persuasion tactics to influence political opinion without disclosing intent. How would you classify this against a policy that permits factual discussion but prohibits covert persuasion?”

  2. 2.Hallucination detection accuracy

    Distinguishing genuine hallucinations from rare but plausible outputs is critical; false positives undermine training, while missed hallucinations create safety gaps.

    For example: “An AI response cites a scientific study with realistic formatting and details, but you suspect it's fabricated. You cannot look it up in the evaluation window. What approach would you take to assess whether this is a hallucination?”

  3. 3.Gray-area judgment calls

    Many safety scenarios lack clear-cut answers; evaluators must apply judgment consistently while acknowledging genuine ambiguity in policy-sensitive topics.

    For example: “A response contains technically accurate information about a cybersecurity vulnerability but could potentially enable malicious use. The policy explicitly permits educational security content. How would you evaluate this?”

  4. 4.Reasoning failure analysis

    Identifying when models reach correct conclusions through flawed logic or incorrect intermediate steps helps teams improve model reasoning quality, not just final answers.

    For example: “An AI response arrives at the correct policy recommendation but uses several questionable logical leaps and mischaracterizes a prior precedent in the reasoning. How would you structure feedback?”

  5. 5.Rubric consistency application

    Consistent rubric application ensures training data quality and prevents bias; inconsistency can introduce systematic errors that models learn from.

    For example: “You've evaluated 50 responses using a safety rubric. You notice your severity ratings for ambiguous cases drifted from strict to lenient. How would you recalibrate consistency?”

A task you may get

Evaluate a set of AI responses across varied risk domains, identifying policy violations, hallucinations, and reasoning failures, then providing structured feedback.

How to prepare

  • Review current AI safety policies, common violation patterns, and evaluation rubric frameworks
  • Study types of AI hallucinations and reasoning failures to develop recognition skills
  • Research policy-sensitive topics like misinformation, persuasion, biosecurity to build domain context
  • Practice giving structured, actionable feedback that identifies problems and explains significance

The facts

Pay
$60–70/hr
Hours
Hourly, 40 hours a week
Where
Remote
Open to
USA, DNK, EST, FIN, ISL, IRL, LVA, LTU, NOR, SWE, AUT, BEL, FRA, DEU, LIE, LUX, MCO, NLD, CHE, GBR, ALB, BIH, HRV, GRC, ITA, XKX, MLT, MKD, PRT, SMR, SRB, SVN, ESP, BGR, CZE, HUN, MDA, POL, ROU, SVK, USA, DNK, EST, FIN, ISL, IRL, LVA, LTU, NOR, SWE, AUT, BEL, FRA, DEU, LIE, LUX, MCO, NLD, CHE, GBR, ALB, BIH, HRV, GRC, ITA, XKX, MLT, MKD, PRT, SMR, SRB, SVN, ESP, BGR, CZE, HUN, MDA, POL, ROU, SVK
Field
Life, Physical, and Social Science
Project name
AIUC
Posted
7/16/2026
Places left
4

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.