$29–45/hr · Mercor · Hourly, 40 hours a week
You red-team AI models with adversarial inputs to uncover vulnerabilities and generate datasets that make AI systems safer for production use.
What you would do
- Red team conversational AI systems with jailbreaks, prompt injections, misuse cases, bias exploitation attempts, and multi-turn manipulation tactics
- Generate high-quality human-authored adversarial data: annotate model failures, classify the type and severity of each vulnerability, and identify patterns
- Follow structured taxonomies, benchmarks, and testing playbooks to ensure consistent methodology and reproducible results across test cases
- Produce detailed attack datasets and vulnerability reports with concrete examples that engineering teams can review and act on
- Communicate technical findings to both AI researchers and non-technical stakeholders, explaining why specific vulnerabilities matter and what customer impact they pose
Who they want
- Prior red-teaming, adversarial-ML, or cybersecurity experience where you probed systems for vulnerabilities and documented failures
- Curious and adversarial mindset: you instinctively look for system weak points and creative ways to trigger unexpected behavior
- Structured approach to vulnerability research using frameworks, taxonomies, or benchmarks rather than random hacking
- Fluent English and Portuguese (global, excluding Brazilian Portuguese); native-level communication in both languages
- Ability to work across projects and customers, context-switch between adversarial tasks, and thrive in ambiguity; psychological resilience for sustained work on harmful scenarios
Main skills
What the interview asks about
1.Multi-turn jailbreak orchestration
Single-turn safety tests miss sophisticated attacks; understanding how to chain requests across multiple turns is what separates advanced red teamers from entry-level testers.
For example: “You want to get a model to generate harmful content across two turns without triggering guardrails on either turn. Sketch a two-turn attack that uses context-switching or role-play to evade safety detection.”
2.Vulnerability pattern recognition
Effective red teamers find not just one failure but the class of failures; categorizing patterns helps teams prioritize remediation.
For example: “You have found three different prompts that make a model violate its safety policy in similar ways. How would you analyze them to identify the root vulnerability rather than treating each as isolated?”
3.Structured methodology discipline
Ad-hoc hacking does not scale; reproducible testing uses consistent frameworks so findings generalize and teams can verify fixes.
For example: “Design a taxonomy of five vulnerability types relevant to conversational AI safety (e.g., jailbreak, reward hacking, context confusion). For each, write one concrete test case that clearly demonstrates the class.”
4.Communicating risk to non-specialists
Engineers need clear, reproducible attack cases; non-technical stakeholders need to understand business and user-safety implications without deep ML knowledge.
For example: “You discovered a jailbreak in an HR chatbot that extracts confidential salary data across two turns. Write a two-paragraph explanation of the vulnerability and its impact for (a) an ML engineer and (b) a company security officer.”
5.Socio-technical and bias-oriented probing
Sophisticated attacks exploit social psychology, cultural bias, and multi-step manipulation that vanilla prompt injection tests miss.
For example: “A model has been trained to give investment advice. Design an attack that uses cultural assumptions or conversation dynamics to trick it into giving harmful financial guidance while appearing reasonable.”
A task you may get
Create a structured red-team test suite with five adversarial prompts across categories: benign, jailbreak, prompt injection, reasoning exploit, and edge case.
How to prepare
- Study one published jailbreak or prompt injection technique in detail (e.g., DAN, role-play persona jailbreaks) and experiment with variations to understand why it works
- Review a taxonomy or benchmark for adversarial testing (e.g., from academic literature) and adapt it to your language and domain
- Practice writing clear, non-technical vulnerability descriptions for someone unfamiliar with AI safety research
- Document two conversational manipulation tactics (e.g., goal-shifting, context confusion) and sketch how an AI system could be vulnerable to each
The facts
- Pay
- $29–45/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Miscellaneous
- Project name
- Neon
- Posted
- 7/30/2026
- Places left
- 20
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.