$16–22/hr · Mercor · Hourly, 40 hours a week
An AI safety specialist who probes conversational models with adversarial inputs to surface vulnerabilities and generate training data.
What you would do
- Design and execute red team attacks on AI models targeting jailbreaks, prompt injections, misuse cases, and bias exploitation
- Generate high-quality labeled data that captures failure modes, classifying vulnerabilities and flagging systemic risks
- Apply published taxonomies and testing playbooks to ensure consistency and rigor across all assessment work
- Document reproducible attack scenarios, datasets, and findings so downstream teams can strengthen models
Who they want
- Fluent in both English and Punjabi with native or near-native command of nuance and cultural context
- Strong judgment about language: ability to explain why an AI response succeeds or fails, with clear reasoning
- Rigorous attention to detail: you catch inconsistencies and logical gaps that less careful reviewers miss
- Experience communicating complex findings to both technical and non-technical stakeholders in writing
Main skills
What the interview asks about
1.Adversarial prompt design
Your ability to construct unexpected inputs that bypass safety guardrails determines how well your testing uncovers real vulnerabilities.
For example: “Walk me through how you would design a multi-turn conversation that gradually escalates requests to push a model toward generating harmful content, and how you'd document each turning point.”
2.Vulnerability categorization
Distinguishing between critical failures and benign edge cases is essential so clients prioritize fixes effectively and don't waste resources on low-risk findings.
For example: “You're testing a model trained to refuse medical advice. It declines to diagnose symptoms but suggests three over-the-counter remedies for someone describing clear pneumonia signs. How would you classify this failure and explain why it matters?”
3.Bias detection in context
Bias often emerges through subtle patterns across many queries rather than a single failure, and your ability to surface these trends shapes how clients understand their model's blind spots.
For example: “Across 50 job recommendation queries, the model ranks males higher for engineering roles, females higher for HR. How would you present this pattern? What follow-up testing would you recommend?”
4.Documentation for technical audiences
Your written explanations are how engineers understand which inputs triggered failures and how to reproduce them, directly determining whether fixes actually address root causes.
For example: “Describe how you would document a jailbreak attempt that required 8 steps to succeed, including the exact model responses at each stage, so another researcher can verify your work and an engineer can trace the vulnerability.”
5.Cross-cultural language judgment
Your bilingual fluency lets you catch culturally contextual biases and meaning shifts that monolingual testers miss, which is essential for models serving diverse user bases.
For example: “A model trained primarily on English data sometimes generates responses in Punjabi that are technically grammatical but culturally inappropriate or offensive. How would you systematically surface these issues in your testing report?”
A task you may get
Conduct a simulated red team assessment of a conversational AI model, crafting adversarial prompts targeting three distinct vulnerability classes, documenting each attempt with the model's response and your classification of the failure mode.
How to prepare
- Study published AI safety benchmarks and red team frameworks to understand standard taxonomies and testing methodologies
- Prepare examples from your experience where you identified subtle errors or misalignments in language or logic that others overlooked
- Review case studies of real AI failures in bias, hallucination, or harmful output generation to ground your understanding of stakes
The facts
- Pay
- $16–22/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Project name
- Neon
- Posted
- 6/4/2026
- Places left
- 1
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.