Training Turk

AI Safety Experts — English & Norwegian

$48–62/hr · Mercor · Hourly, 40 hours a week

Conduct AI red teaming in English and Norwegian to identify vulnerabilities and generate safety training data.

What you would do

  • Design adversarial scenarios targeting model biases, misuse potential, and manipulation vulnerabilities
  • Apply jailbreak and prompt injection techniques systematically following documented playbooks
  • Create multi-turn conversations that probe model consistency and value alignment
  • Generate annotated datasets classifying failures by vulnerability type and severity
  • Document reproducible attacks and their implications for real-world deployment

Who they want

  • Prior red teaming, adversarial ML, or cybersecurity experience
  • Native fluency in both English and Norwegian
  • Curiosity and systematic approach to finding model breaking points
  • Skill at articulating vulnerabilities and security implications to both technical experts and general audiences
  • Adaptability across different model types and customer contexts

Main skills

Red teamingAdversarial input designJailbreak techniques

What the interview asks about

  1. 1.Novel attack generation

    Existing playbooks become stale; your ability to innovate determines long-term data value.

    For example: “A model has been hardened against direct jailbreaks. Describe a multi-turn conversation approach that exploits different vulnerabilities in sequence.”

  2. 2.Cross-lingual bias detection

    Biases manifest differently in Norwegian versus English; bilingual expertise matters.

    For example: “Design a scenario to test whether a model exhibits different gender stereotypes when responding in Norwegian versus English.”

  3. 3.Distinguishing benign from harmful ambiguity

    Not every surprising output is a vulnerability; your judgment guides training data quality.

    For example: “A model gives an odd response to a Norwegian idiom. How would you assess whether this is a concerning gap or expected behavior?”

  4. 4.Documenting reproducibility

    Training data must contain actionable, reproducible findings.

    For example: “You found an attack but it works inconsistently. How would you modify your approach to make it reliable and document it?”

A task you may get

Conduct red teaming on a provided prompt: find 3-4 distinct vulnerabilities, document each with reproducible steps and risk assessment.

How to prepare

  • Study existing red teaming frameworks and taxonomies for classifying failures
  • Review examples of jailbreaks, prompt injections, and bias exploitations
  • Think about cultural or linguistic differences that might create vulnerabilities in Norwegian vs English
  • Prepare 2-3 original attack ideas you could articulate step-by-step

The facts

Pay
$48–62/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Project name
Neon
Posted
7/30/2026
Places left
20

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.