$48–62/hr · Mercor · Hourly, 40 hours a week
You adversarially test conversational AI models to uncover safety vulnerabilities and generate training data for AI safety improvement.
What you would do
- Design and execute red teaming attacks: jailbreaks, prompt injections, multi-turn manipulation, and bias exploitation strategies
- Annotate model failures, classify vulnerabilities by type and severity, and flag systemic risks in AI responses
- Apply structured frameworks and taxonomies to conduct consistent testing across different models and scenarios
- Produce detailed, reproducible attack documentation and datasets that customers can validate and remediate
- Communicate findings clearly to both technical teams and non-technical stakeholders about risk and impact
Who they want
- Native fluency in English and Danish required for this role
- Prior experience with red teaming, adversarial AI testing, cybersecurity, or socio-technical risk assessment
- Curious, adversarial mindset with instinct to push systems to failure points systematically
- Structured approach using frameworks, not random attacks; ability to document findings clearly
- Comfortable discussing sensitive topics like bias, misinformation, and harmful content with clear ethical guidelines
Main skills
What the interview asks about
1.Jailbreak design and execution
Demonstrates understanding of model constraints and ability to design multi-step attacks that surface inconsistencies in safety training.
For example: “Design a three-turn conversation that gradually escalates from benign requests to increasingly harmful content. How would you measure when the model has failed?”
2.Bias discovery in multilingual contexts
Tests whether the candidate recognizes that bias manifests differently across languages and cultural contexts, essential for Danish language testing.
For example: “You're testing whether a model treats Danish versus English requests for sensitive topics differently. What attack patterns would you use to detect this disparity?”
3.Taxonomy and framework application
Shows ability to work systematically rather than chaotically, which separates effective red teamers from those producing noisy, irreproducible results.
For example: “Map a vulnerability you discovered to both the OWASP AI taxonomy and your project's internal taxonomy. Explain why multiple categorizations matter for remediation.”
4.Documentation for customer action
Reveals whether findings are communicated as isolated examples or as reproducible, systematized attacks customers can actually fix.
For example: “You discovered that the model fails when given role-play scenarios framed in Danish versus English. Write a report summarizing the vulnerability and reproducible steps.”
5.Prioritizing vulnerabilities by impact
Tests judgment about which failures matter most and why, ensuring effort focuses on genuine risks rather than theoretical edge cases.
For example: “You've found 20 potential jailbreaks this week. Your customer can fix 3 before launch. Which three do you prioritize and why?”
A task you may get
Design a three-part attack sequence targeting a conversational AI's handling of misinformation, then document one successful failure case with steps a customer engineer could reproduce and fix.
How to prepare
- Review published frameworks for AI safety testing (OWASP, NIST, or similar) to understand structured red teaming methodologies
- Study at least two research papers on adversarial attacks in conversational AI to learn sophisticated attack patterns
- Practice writing technical findings so non-engineers can understand vulnerability severity and remediation options
The facts
- Pay
- $48–62/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Data Analysis
- Project name
- Neon
- Posted
- 7/30/2026
- Places left
- 20
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.