$17–25/hr · Mercor · Hourly, 40 hours a week
Red team AI conversational systems by designing and executing adversarial attacks to surface vulnerabilities.
What you would do
- Design targeted adversarial prompts to expose AI model weaknesses and security gaps
- Classify and document vulnerability findings using standardized taxonomies and frameworks
- Generate reproducible test datasets showing attack patterns and failure modes
- Evaluate multi-turn conversation scenarios to identify cumulative manipulation risks
- Communicate technical findings and business implications to both engineering and stakeholder audiences
Who they want
- Prior experience in AI red teaming, security testing, or adversarial machine learning
- Structured thinking: ability to apply frameworks rather than use random hacking approaches
- Clear writing and verbal communication skills for diverse audiences
- Curiosity about system failure modes and willingness to probe breaking points
- Optional: background in cybersecurity, jailbreak research, RLHF attacks, or socio-technical risk analysis
Main skills
What the interview asks about
1.Jailbreak technique design
Interviewers need to verify you can systematically probe AI guard rails and identify which attack patterns work against different model architectures.
For example: “Describe three distinct jailbreak approaches you'd test against a customer service chatbot, and how you'd document whether each technique bypassed its safety layer.”
2.Failure taxonomy application
Demonstrating a structured approach to classification shows you can generate consistent, comparable datasets rather than subjective assessments.
For example: “Given AI outputs from bias-exploitation probes on a loan approval model, how would you classify and prioritize which failures represent the biggest fairness risks?”
3.Reproducible attack documentation
Customers trust your data only if they can replicate your findings; this tests your ability to write technical procedures others can follow.
For example: “Walk through how you'd document a successful prompt injection attack on a code-generation model so another security engineer could reproduce it exactly.”
4.Multi-turn manipulation strategy
Real-world risks come from sustained conversations, not single prompts; this tests whether you understand context accumulation and behavioral drift.
For example: “Design a five-turn conversation sequence that gradually shifts an AI assistant's output toward harmful content, explaining how each turn builds on prior responses.”
5.Stakeholder risk communication
Your findings only matter if decision-makers grasp both the technical details and the business impact; this tests translation between audiences.
For example: “A model's bias vulnerability could affect loan approvals for 20% of applicants from a particular demographic. How would you present this risk to both the engineering team and business leadership?”
6.Framework consistency under pressure
Testing speed and volume can tempt shortcuts; this checks whether you maintain rigor when workload increases.
For example: “You have 15 hours to test 40 new attack scenarios. How would you ensure your taxonomy labels remain consistent without rushing classifications?”
A task you may get
Design a five-prompt red team strategy targeting a customer service bot, predict what vulnerabilities each might expose, and write a one-page report showing how results would be documented for a customer review.
How to prepare
- Study OWASP adversarial ML frameworks and standardized taxonomies for AI safety testing
- Collect 3-5 examples of AI jailbreaks or prompt injections from published research and practice classifying them
- Prepare two case studies from your own security testing showing how you documented and communicated findings
- Review the difference between ad-hoc fuzzing and structured taxonomy-based testing
The facts
- Pay
- $17–25/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Project name
- Neon
- Posted
- 7/30/2026
- Places left
- 20
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.