$16–22/hr · Mercor · Hourly, 40 hours a week
You conduct systematic adversarial testing against AI models to uncover safety weaknesses and generate structured data enabling safer AI deployment.
What you would do
- Design and execute adversarial prompts targeting model vulnerabilities using injection, jailbreak, and manipulation techniques
- Analyze AI outputs for accuracy issues, logical inconsistency, and cultural misalignment across English and Marathi contexts
- Classify each failure using established taxonomies to create comparable, actionable vulnerability records
- Synthesize individual test results into pattern analysis that reveals systemic model weaknesses
Who they want
- Native-level fluency in English and Marathi required for context-aware evaluation
- Sharp linguistic judgment distinguishing between legitimate and inappropriate AI responses
- Methodical approach to applying standards and benchmarks without deviation
- Clear communication ability explaining technical vulnerabilities to non-technical stakeholders
Main skills
What the interview asks about
1.Unconventional testing approaches
Safety gaps often lie in scenarios that don't follow obvious attack patterns; finding these requires combining strategies and learning from failures to iterate.
For example: “Your direct jailbreak attempts failed on a customer's chatbot. Describe how you would modify your approach to bypass its defenses using multi-turn conversation or contextual misdirection.”
2.Marathi-specific vulnerability detection
Languages encode cultural context differently; an appropriate response in English might be offensive in Marathi or vice versa, creating blind spots for monolingual testers.
For example: “Testing reveals an AI gives different financial advice based on whether the question is asked in English or Marathi. How would you structure your findings to help the team fix this?”
3.Structured documentation under ambiguity
When a vulnerability doesn't fit standard categories, poor documentation either wastes engineering time or lets similar issues slip into production.
For example: “You found a failure that appears to mix bias, misinformation, and language handling. Walk me through how you'd break this into classifiable components for the taxonomy.”
4.Quality maintenance at scale
As test volume increases, fatigue leads people to truncate documentation, skip edge cases, or apply categories inconsistently, degrading dataset value.
For example: “By hour 8 of testing today, you notice classification decisions are harder. How do you ensure your later test outputs meet the same standard as your first ones?”
5.Pattern extraction from test results
Isolated failures matter less than recognizing trends that point to architectural issues, enabling focused, efficient fixes across multiple scenarios.
For example: “Across 15 test cases involving relationship advice, the model consistently fails to maintain empathy when discussing cultural differences. How would you investigate and report this pattern?”
A task you may get
Given a Marathi-English code-switched chatbot sample, identify one jailbreak vulnerability and one bias vulnerability, then classify each using a provided framework with detailed reasoning.
How to prepare
- Study adversarial ML attack patterns: jailbreaks, prompt injection, and multi-step manipulation in both languages
- Research cultural contexts and communication norms specific to Marathi to identify contextual vulnerabilities
- Review data classification taxonomies for bias, misinformation, and safety domains
- Practice articulating technical findings for audiences ranging from engineers to business stakeholders
The facts
- Pay
- $16–22/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Project name
- Neon
- Posted
- 6/3/2026
- Places left
- 7
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.