Training Turk

AI Safety Experts — English & Swedish

$48–62/hr · Mercor · Hourly, 40 hours a week

You red team AI models by crafting adversarial inputs to uncover safety vulnerabilities and generate data that makes models more robust.

What you would do

  • Design and execute red team attacks targeting jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation scenarios
  • Generate high-quality attack data by crafting adversarial examples and documenting model failures with structured classification
  • Annotate vulnerabilities using consistent taxonomies to keep red teaming reproducible and measurable
  • Develop reproducible attack cases with clear documentation and exploitation paths for teams to learn from
  • Assess vulnerabilities by severity, impact, and real-world likelihood; communicate findings to diverse stakeholders

Who they want

  • Prior red teaming or adversarial experience in AI testing, cybersecurity, or penetration testing
  • Curiosity and adversarial mindset: you instinctively probe systems and push boundaries to find breaking points
  • Structured thinking: you use frameworks, playbooks, and taxonomies rather than random testing
  • Clear communication skills to explain vulnerabilities, severity, and remediation to engineers and executives
  • Adaptability across projects with ability to shift domains and threat models quickly

Main skills

Adversarial attack designJailbreak techniquesPrompt injection methods

What the interview asks about

  1. 1.Jailbreak and prompt injection design

    Creative adversarial prompts uncover model vulnerabilities that standard inputs miss; your toolkit and reasoning reveal your depth.

    For example: “You're red teaming a customer support chatbot that should refuse requests to reset passwords without verification. Walk me through 3-4 different jailbreak or prompt injection approaches you'd try, explaining why each targets a different model weakness.”

  2. 2.Systematic attack taxonomy and classification

    Ad-hoc testing is unpredictable; frameworks and taxonomies make attacks reproducible and help teams defend systematically.

    For example: “You've uncovered 10 different jailbreaks against a model. How would you classify them into a taxonomy that customers could use to prioritize fixes and measure progress? What categories matter most?”

  3. 3.Multi-turn and behavioral exploitation

    Single-turn attacks are obvious; sophisticated adversaries use multi-turn sequences and behavioral shifts to gradually shift model reasoning.

    For example: “Describe a multi-turn attack scenario where you'd gradually shift a model's stated policies or safety constraints. Walk me through the progression and what you'd document to make it reproducible.”

  4. 4.Documenting and communicating vulnerability impact

    Engineers need to understand root cause; executives need to understand business risk; your documentation bridges both.

    For example: “You've found a prompt injection that extracts the model's internal system prompt. How would you document this for an engineer investigating the root cause, and separately, what would you tell a customer CEO about the business impact?”

  5. 5.Bias and socio-technical attack vectors

    Safety isn't just technical; attacks that exploit bias, social manipulation, or downstream harm require different thinking.

    For example: “A model generates plausible-sounding medical advice for a rare disease that's actually dangerous. Is this a jailbreak, a bias, or a training data issue? How would you classify it and what would you recommend for testing and defense?”

A task you may get

Identify 3-4 red team attack vectors for a conversational AI scenario. Classify using your taxonomy. Write one reproducible attack with steps, expected failure, and business impact.

How to prepare

  • Review a recent AI safety incident or vulnerability disclosure and analyze the attack chain: how did the attacker probe, what did they find, how would you have detected it earlier?
  • Study 2-3 OWASP or security testing frameworks and think through how to adapt them for conversational AI
  • Prepare a concrete example of a multi-turn attack or manipulation chain you've executed or researched, describing the progression and how you'd document it for reproducibility
  • Reflect on how you'd explain a technical vulnerability to a non-technical executive: what's the impact, who's at risk, and what should they fund to fix it?

The facts

Pay
$48–62/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Project name
Neon
Posted
7/30/2026
Places left
20

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.