$48–62/hr · Mercor · Hourly, 40 hours a week
Conduct AI red teaming in English and Norwegian to identify vulnerabilities and generate safety training data.
What you would do
- Design adversarial scenarios targeting model biases, misuse potential, and manipulation vulnerabilities
- Apply jailbreak and prompt injection techniques systematically following documented playbooks
- Create multi-turn conversations that probe model consistency and value alignment
- Generate annotated datasets classifying failures by vulnerability type and severity
- Document reproducible attacks and their implications for real-world deployment
Who they want
- Prior red teaming, adversarial ML, or cybersecurity experience
- Native fluency in both English and Norwegian
- Curiosity and systematic approach to finding model breaking points
- Skill at articulating vulnerabilities and security implications to both technical experts and general audiences
- Adaptability across different model types and customer contexts
Main skills
What the interview asks about
1.Novel attack generation
Existing playbooks become stale; your ability to innovate determines long-term data value.
For example: “A model has been hardened against direct jailbreaks. Describe a multi-turn conversation approach that exploits different vulnerabilities in sequence.”
2.Cross-lingual bias detection
Biases manifest differently in Norwegian versus English; bilingual expertise matters.
For example: “Design a scenario to test whether a model exhibits different gender stereotypes when responding in Norwegian versus English.”
3.Distinguishing benign from harmful ambiguity
Not every surprising output is a vulnerability; your judgment guides training data quality.
For example: “A model gives an odd response to a Norwegian idiom. How would you assess whether this is a concerning gap or expected behavior?”
4.Documenting reproducibility
Training data must contain actionable, reproducible findings.
For example: “You found an attack but it works inconsistently. How would you modify your approach to make it reliable and document it?”
A task you may get
Conduct red teaming on a provided prompt: find 3-4 distinct vulnerabilities, document each with reproducible steps and risk assessment.
How to prepare
- Study existing red teaming frameworks and taxonomies for classifying failures
- Review examples of jailbreaks, prompt injections, and bias exploitations
- Think about cultural or linguistic differences that might create vulnerabilities in Norwegian vs English
- Prepare 2-3 original attack ideas you could articulate step-by-step
The facts
- Pay
- $48–62/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Project name
- Neon
- Posted
- 7/30/2026
- Places left
- 20
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.