Training Turk

AI Safety Experts — English & Bengali

$16–22/hr · Mercor · Hourly, 40 hours a week

You red-team conversational AI systems by finding vulnerabilities through adversarial testing and generating safety datasets.

What you would do

  • Craft adversarial prompts to probe AI models including jailbreaks, prompt injections, misuse cases, and multi-turn manipulation
  • Annotate and classify AI failures using defined taxonomies to categorize vulnerability types, severity, and systemic patterns
  • Evaluate AI outputs for accuracy, completeness, appropriateness, and potential harms across bias and misinformation topics
  • Generate high-quality datasets of failure cases and attack scenarios that customers use to improve model robustness
  • Document findings and reproducible attack cases so clients can validate vulnerabilities and verify fixes

Who they want

  • Native fluency in English and Bengali required for all work
  • Strong language judgment: assess whether AI responses are accurate, complete, appropriate, and free of bias
  • Rigorous and detail-oriented: notice subtle inconsistencies and errors that casual review misses
  • Structured thinking: work consistently to defined guidelines and quality standards, not ad hoc approaches
  • Adaptable mindset: comfort with diverse projects, task types, and shifting priorities as customer needs evolve

Main skills

Red team testingAdversarial promptingJailbreak development

What the interview asks about

  1. 1.Finding bias in AI outputs

    Bias often hides in subtle language and unstated assumptions; this requires cultural and linguistic judgment to catch.

    For example: “An AI summarizes hiring criteria for two roles but subtly emphasizes 'leadership' for one and 'technical excellence' for the other. How would you frame this as a vulnerability?”

  2. 2.Crafting multi-turn jailbreaks

    Single-prompt attacks are easy; the complexity comes from sustained manipulation over multiple turns revealing strategic thinking.

    For example: “An AI refuses harmful requests. Your first jailbreak fails. Describe three escalating approaches and how each pivots the attack.”

  3. 3.Distinguishing vulnerabilities from quirks

    Not every unexpected output is a vulnerability; you need judgment to flag what matters for safety versus model idiosyncrasies.

    For example: “An AI answers medications questions with less detail in one language pair than English, with more caveats. Is this a gap, bias, or expected behavior?”

  4. 4.Structured taxonomy application

    Consistency and reproducibility matter; interviewers verify you apply taxonomies disciplined rather than inventing categories ad hoc.

    For example: “You identify 15 failures in 100 responses. Your taxonomy has 7 types. Three failures don't fit neatly. How do you handle classification?”

  5. 5.Managing exposure to harmful content

    Red teaming exposes you to harmful content intentionally; demonstrating professional approach and resource use shows maturity.

    For example: “After 2 hours probing misinformation, you feel fatigued from exposure to harmful narratives. What's your next step?”

  6. 6.Evaluating content accuracy and appropriateness

    Assessing whether AI responses are truthful, balanced, and suitable for audiences requires deep judgment about language and context.

    For example: “An AI describes a medical treatment with accurate facts but omits critical risks. Is this a safety failure, and how would you report it?”

A task you may get

Design a multi-turn attack on a conversational AI system that attempts a jailbreak or bias exploit. Document each turn, your hypothesis, and adjustments based on responses.

How to prepare

  • Study published jailbreak techniques and prompt injection papers to learn common patterns and emerging methods
  • Research AI safety taxonomies and vulnerability classification schemes used by major AI providers
  • Practice close reading: analyze AI outputs for biases, incomplete reasoning, and unstated assumptions in diverse topics
  • Reflect on your language judgment: what makes a response to sensitive topics appropriate versus harmful or biased?

The facts

Pay
$16–22/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Miscellaneous
Project name
Neon
Posted
6/4/2026
Places left
14

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.