$16–22/hr · Mercor · Hourly, 40 hours a week
An AI Safety Expert red-teams AI systems to identify vulnerabilities and generate training data that makes models safer.
What you would do
- Red team conversational AI models with adversarial inputs to surface jailbreaks, prompt injections, bias exploitation, and manipulation tactics
- Analyze and annotate AI failures, classifying vulnerability types and flagging systemic risks in model behavior
- Follow established taxonomies and playbooks to ensure testing consistency and structured documentation across projects
- Generate reproducible reports and datasets that customers use to understand vulnerabilities and improve system robustness
- Evaluate AI outputs for accuracy, completeness, bias, and appropriateness across sensitive content domains
Who they want
- Native fluency in English and Assamese (both languages required for this role)
- Strong judgment about language and content with ability to identify whether AI responses are accurate and appropriate
- Rigorous attention to detail, noticing subtle errors and inconsistencies in AI outputs that others overlook
- Structured approach to work, applying guidelines and quality standards consistently rather than ad hoc
- Clear communication skills for explaining findings and reasoning to technical and non-technical audiences
What the interview asks about
1.Adversarial thinking and attack design
Red teamers must think like attackers to find novel vulnerabilities; interviews probe whether you can generate unconventional attack patterns.
For example: “You're testing an AI chatbot trained with RLHF to refuse harmful requests. Design 3 different attack approaches that use different techniques: one exploiting role-play, one using technical language, one using emotional appeals. Explain why each might succeed.”
2.Language judgment in context
Bilingual fluency alone is insufficient; evaluators must judge whether AI responses are culturally appropriate, factually accurate, and complete in both languages.
For example: “An AI responds to an Assamese query about traditional medicine in English instead of Assamese, then provides partial information missing key context. How would you evaluate this failure, and what would you document?”
3.Distinguishing intentional safety vs. genuine failure
AI systems may refuse requests intentionally (good) or accidentally refuse reasonable requests (bad); testers must distinguish between these using judgment.
For example: “An AI refuses a straightforward factual question about public health policy in Assamese because it detects the word 'pandemic.' Is this an over-trigger, under-trigger, or appropriate safety response? How do you evaluate it?”
4.Structured taxonomy application
Consistent classification is essential so vulnerability data is usable; interviews probe whether you can apply structured frameworks reliably across edge cases.
For example: “You're using a 5-category taxonomy for bias failures. You find a response that exhibits two types of bias simultaneously, plus factual inaccuracy. How do you classify this and document it so researchers can interpret your judgment?”
5.Reproducible documentation
Red team findings must enable customer engineering teams to replicate and fix vulnerabilities; vague reports waste effort on reproduction.
For example: “You found a jailbreak that works 70% of the time with a multi-turn prompt sequence. What specific details would you include in your report to make it reproducible, and what would make reproduction fail?”
6.Sensitive content handling
This role involves reviewing potentially harmful content for testing purposes; interviews probe judgment about what counts as harmful and how to document it safely.
For example: “You're testing an AI's refusals on misinformation. You need to generate realistic misinformation examples to test against, but don't want to create genuinely harmful content. How do you navigate this boundary?”
A task you may get
Given 5 conversational AI queries in English and Assamese spanning benign, sensitive, and harmful topics, identify vulnerabilities, classify them using a provided taxonomy, and write a report with reproducible attack sequences for each identified failure.
How to prepare
- Study common jailbreak techniques and prompt injection methods used against LLMs to understand attack surface
- Review AI safety taxonomies and classification schemes to understand how vulnerabilities are typically categorized
- Practice analyzing bilingual AI responses for cultural appropriateness, factual accuracy, and subtle bias in both languages
- Research documented AI failures and red team datasets to see what makes findings reproducible and actionable for developers
The facts
- Pay
- $16–22/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Project name
- Neon
- Posted
- 6/4/2026
- Places left
- 6
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.