$16–22/hr · Mercor · Hourly, 40 hours a week
A red-teaming specialist in English and Malayalam who crafts adversarial inputs for conversational AI and evaluates whether models resist jailbreaks and biases safely.
What you would do
- Design adversarial prompts targeting jailbreaks, prompt injections, bias, and multi-turn manipulation
- Evaluate AI responses for accuracy, completeness, and safety alignment
- Classify vulnerabilities and failure modes using structured taxonomies
- Annotate datasets and flag systemic risks detected across multiple test cases
- Document findings and attack sequences clearly for technical teams
Who they want
- Fluent English and native Malayalam language skills required
- Strong judgment about language, content, and AI appropriateness
- Attention to detail and ability to spot subtle errors and inconsistencies
- Rigorous approach to quality standards and taxonomy compliance
- Communication skill to explain findings to both technical and non-technical audiences
Main skills
What the interview asks about
1.Jailbreak prompt design
Effective jailbreaks reveal real model vulnerabilities; poorly designed attacks miss actual weaknesses or trigger obvious safe guards.
For example: “A model is trained to refuse code generation for hacking tools. Design 3 different jailbreak approaches you'd test, and explain what each reveals about the model's robustness.”
2.Bias detection across language
Biases in multilingual models often hide in cultural assumptions or translation artifacts; detecting these requires language fluency.
For example: “Testing bias in a Malayalam conversational AI, how would you design prompts about gender roles, caste, or religion that expose whether the model absorbs training data biases?”
3.Multi-turn manipulation sequence
Single-turn attacks are easier to defend; sophisticated attacks accumulate context or state across turns to defeat guards.
For example: “Design a 4-turn conversation where each turn builds on the previous one to gradually push a model toward generating harmful content. What vulnerabilities does this attack expose?”
4.Taxonomy compliance and annotation
Consistent classification enables research teams to identify patterns; ad-hoc labeling produces unusable data.
For example: “You encounter a response that mixes accurate information with subtle misinformation. Which failure category would you assign, and what label would you apply to train future detectors?”
5.Systemic risk identification
Individual failures matter less than patterns; spotting that a model consistently fails on a specific attack type is actionable research.
For example: “Across 20 test cases, you notice the model confidently generates incorrect reasoning for mathematical word problems. How would you document this systemic risk to guide model improvement?”
A task you may get
Design 3 adversarial prompts (one jailbreak, one bias test, one multi-turn manipulation) for a conversational AI, evaluate simulated responses, and classify failures using a provided taxonomy.
How to prepare
- Study common jailbreak patterns and adversarial NLP techniques
- Reflect on bias patterns in large language models trained on diverse data
- Practice explaining technical AI concepts to non-specialists clearly
- Review structured taxonomies for AI failure modes and vulnerability classification
The facts
- Pay
- $16–22/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Miscellaneous
- Project name
- Neon
- Posted
- 6/4/2026
- Places left
- 1
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.