$48–62/hr · Mercor · Hourly, 40 hours a week
You red team AI models by crafting adversarial inputs to uncover safety vulnerabilities and generate data that makes models more robust.
What you would do
- Design and execute red team attacks targeting jailbreaks, prompt injection, bias exploitation, and multi-turn manipulation scenarios
- Generate high-quality attack data by crafting adversarial examples and documenting model failures with structured classification
- Annotate vulnerabilities using consistent taxonomies to keep red teaming reproducible and measurable
- Develop reproducible attack cases with clear documentation and exploitation paths for teams to learn from
- Assess vulnerabilities by severity, impact, and real-world likelihood; communicate findings to diverse stakeholders
Who they want
- Prior red teaming or adversarial experience in AI testing, cybersecurity, or penetration testing
- Curiosity and adversarial mindset: you instinctively probe systems and push boundaries to find breaking points
- Structured thinking: you use frameworks, playbooks, and taxonomies rather than random testing
- Clear communication skills to explain vulnerabilities, severity, and remediation to engineers and executives
- Adaptability across projects with ability to shift domains and threat models quickly
Main skills
What the interview asks about
1.Jailbreak and prompt injection design
Creative adversarial prompts uncover model vulnerabilities that standard inputs miss; your toolkit and reasoning reveal your depth.
For example: “You're red teaming a customer support chatbot that should refuse requests to reset passwords without verification. Walk me through 3-4 different jailbreak or prompt injection approaches you'd try, explaining why each targets a different model weakness.”
2.Systematic attack taxonomy and classification
Ad-hoc testing is unpredictable; frameworks and taxonomies make attacks reproducible and help teams defend systematically.
For example: “You've uncovered 10 different jailbreaks against a model. How would you classify them into a taxonomy that customers could use to prioritize fixes and measure progress? What categories matter most?”
3.Multi-turn and behavioral exploitation
Single-turn attacks are obvious; sophisticated adversaries use multi-turn sequences and behavioral shifts to gradually shift model reasoning.
For example: “Describe a multi-turn attack scenario where you'd gradually shift a model's stated policies or safety constraints. Walk me through the progression and what you'd document to make it reproducible.”
4.Documenting and communicating vulnerability impact
Engineers need to understand root cause; executives need to understand business risk; your documentation bridges both.
For example: “You've found a prompt injection that extracts the model's internal system prompt. How would you document this for an engineer investigating the root cause, and separately, what would you tell a customer CEO about the business impact?”
5.Bias and socio-technical attack vectors
Safety isn't just technical; attacks that exploit bias, social manipulation, or downstream harm require different thinking.
For example: “A model generates plausible-sounding medical advice for a rare disease that's actually dangerous. Is this a jailbreak, a bias, or a training data issue? How would you classify it and what would you recommend for testing and defense?”
A task you may get
Identify 3-4 red team attack vectors for a conversational AI scenario. Classify using your taxonomy. Write one reproducible attack with steps, expected failure, and business impact.
How to prepare
- Review a recent AI safety incident or vulnerability disclosure and analyze the attack chain: how did the attacker probe, what did they find, how would you have detected it earlier?
- Study 2-3 OWASP or security testing frameworks and think through how to adapt them for conversational AI
- Prepare a concrete example of a multi-turn attack or manipulation chain you've executed or researched, describing the progression and how you'd document it for reproducibility
- Reflect on how you'd explain a technical vulnerability to a non-technical executive: what's the impact, who's at risk, and what should they fund to fix it?
The facts
- Pay
- $48–62/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Project name
- Neon
- Posted
- 7/30/2026
- Places left
- 20
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.