$70–84/hr · Mercor · Hourly, 40 hours a week
You design adversarial tests to uncover vulnerabilities in frontier AI systems and collaborate with researchers to improve model safety.
What you would do
- Create challenging multi-turn prompts designed to stress-test model boundaries and expose reasoning failures
- Probe models across sensitive domains (cybersecurity, biosecurity, fraud, political bias, misinformation) to identify unsafe or inconsistent outputs
- Classify vulnerabilities by type, severity, and reproducibility, distinguishing between policy failures and factual errors
- Document detailed findings with examples, context, and recommendations for the research team
- Work closely with safety researchers to validate discoveries and refine testing methodologies
Who they want
- 5+ years in AI safety, red teaming, trust & safety, cybersecurity, investigative journalism, or life sciences
- University-level qualification in computing, cybersecurity, communications, psychology, biology, chemistry, public policy, or related fields
- Expert-level prompt design and adversarial testing skills with experience on frontier systems
- Deep knowledge in at least one sensitive domain like biosecurity, cyber threats, or misinformation
- Strong analytical reasoning and written communication for documenting complex vulnerability findings
Main skills
What the interview asks about
1.Designing layered adversarial prompts
Basic jailbreaks won't uncover the nuanced safety failures that matter. You need to craft multi-step attacks that expose deeper reasoning vulnerabilities.
For example: “A model gives reasonable safety warnings in direct questions about biosynthesis. How would you design a multi-turn prompt sequence to test whether it maintains boundaries under indirect framing or role-play scenarios?”
2.Classifying failure types
Not all model errors are equally concerning. You must distinguish between hallucinations, policy circumvention, and benign knowledge gaps to guide researcher prioritization.
For example: “A model confidently provides a plausible-sounding but false procedure for a sensitive chemistry process. Is this a hallucination, an alignment failure, or a domain knowledge gap? How would you report it?”
3.Evaluating domain-specific safety
Cybersecurity, biosecurity, and financial fraud have distinct threat profiles. You need domain expertise to recognize what outputs are actually dangerous versus what sounds alarming.
For example: “A model is asked about synthetic pathogen creation techniques. What distinctions would you draw between educational information, legitimate research context, and genuinely dangerous guidance?”
4.Communicating vulnerability nuance
Engineers can't fix problems they don't understand. You must explain why a specific output is unsafe, how it was triggered, and what risk it poses.
For example: “You found a pattern where the model's response to political questions shifts based on prompt framing. How would you structure your report to show this is reproducible and explain why consistency matters?”
5.Adapting testing to model constraints
Different models have different architectures and training. Your testing approach must account for what you learn across iterations with the research team.
For example: “After initial testing, researchers tell you the model is harder to jailbreak than expected due to recent training changes. How would you adjust your prompt strategy accordingly?”
A task you may get
Design 3-4 adversarial prompts targeting a frontier AI system in a sensitive domain (biosecurity or political misinformation). Describe the vulnerability each probes for, explain why it's risk-relevant, and predict how different types of models might fail.
How to prepare
- Study recent AI safety incidents and published red-teaming case studies to understand common vulnerability patterns
- Develop deep expertise in one sensitive domain (e.g., biosecurity threats, cybersecurity vulnerabilities, fraud techniques)
- Practice designing multi-turn prompt sequences that layer complexity to bypass surface-level safety guardrails
- Review current alignment techniques (RLHF, SFT, constitutional AI) to understand how models are trained to resist adversarial inputs
The facts
- Pay
- $70–84/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Open to
- USA, DNK, EST, FIN, ISL, IRL, LVA, LTU, NOR, SWE, AUT, BEL, FRA, DEU, LIE, LUX, MCO, NLD, CHE, GBR, ALB, BIH, HRV, GRC, ITA, XKX, MLT, MKD, PRT, SMR, SRB, SVN, ESP, BGR, CZE, HUN, MDA, POL, ROU, SVK, USA, DNK, EST, FIN, ISL, IRL, LVA, LTU, NOR, SWE, AUT, BEL, FRA, DEU, LIE, LUX, MCO, NLD, CHE, GBR, ALB, BIH, HRV, GRC, ITA, XKX, MLT, MKD, PRT, SMR, SRB, SVN, ESP, BGR, CZE, HUN, MDA, POL, ROU, SVK
- Field
- Data Analysis
- Project name
- AIUC
- Posted
- 7/16/2026
- Places left
- 4
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.