Training Turk

AI Safety Experts — English & Vietnamese

$17–25/hr · Mercor · Hourly, 40 hours a week

A bilingual AI safety expert red-teams conversational models to uncover vulnerabilities through structured adversarial testing.

What you would do

  • Craft adversarial inputs and attack scenarios targeting conversational AI models to expose jailbreaks, misuse cases, and bias
  • Execute systematic red team testing following established frameworks and vulnerability taxonomies
  • Document attack methodology, results, and reproducible steps in clear technical reports
  • Classify vulnerabilities by severity and identify systemic patterns that suggest deeper model weaknesses
  • Communicate findings to engineers and non-technical stakeholders in appropriately targeted language

Who they want

  • Fluent English and Vietnamese speaker with native-level capability in both languages
  • Prior red teaming or adversarial AI experience, or proven background in cybersecurity or socio-technical probing
  • Instinct for attacking systems and finding edge cases; curiosity about how things break
  • Structured thinking: you prefer frameworks and benchmarks over ad-hoc hacking
  • Ability to thrive across multiple projects and customer contexts with minimal ramp-up time

Main skills

Adversarial prompt engineeringAI jailbreak techniquesBias exploitation testing

What the interview asks about

  1. 1.Jailbreak construction

    Effective jailbreaks require understanding how AI models interpret instructions and constraints; naive attempts teach the model nothing useful.

    For example: “An AI agent declines financial advice. Identify three jailbreak tactics: character role-play, framing as education, fictional scenarios. Explain which attack pattern each represents and what the model learns from defending.”

  2. 2.Multi-turn manipulation

    Single-turn attacks are often visible; real attacks unfold across multiple model responses with careful context-building to avoid detection.

    For example: “Design a five-turn conversation where you gradually escalate requests to the AI model, starting with safe questions and building to the harmful request by the final turn. What changes between turns that makes the later request seem acceptable in context?”

  3. 3.Bias exploitation specificity

    Bias testing must be concrete and reproducible, not vague; 'make the AI biased' is not a test, but 'apply gender stereotypes to leadership recommendations' is.

    For example: “Test for occupational bias. Describe a structured set of test prompts (not just one) that systematically probe whether the model associates certain professions with specific genders or demographics.”

  4. 4.Vulnerability severity assessment

    Categorizing findings guides customer resource allocation; overstating all issues as critical leads to decision paralysis.

    For example: “Model bypasses safety rules in two scenarios: complex 20-turn attack vs. simple two-turn role-play. Classify each by severity and explain the real-world risk to the customer.”

A task you may get

Design and execute 4 adversarial scenarios against a conversational AI using an 8-category attack taxonomy. Document methodology and classify findings by severity. Evaluated on attack quality and documentation clarity.

How to prepare

  • Study published red team taxonomies (e.g., MLCommons HELM, OpenAI red team papers) to understand structured vulnerability categories
  • Practice designing multi-turn attack scenarios that test specific model blindspots without revealing your intent prematurely
  • Review case studies of real adversarial attacks on AI systems and how they were documented for engineering teams
  • Prepare examples from your own experience of how you explain security risks to non-technical audiences

The facts

Pay
$17–25/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Project name
Neon
Posted
7/30/2026
Places left
20

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.