$48–62/hr · Mercor · Hourly, 40 hours a week
You conduct red team security testing on conversational AI systems to discover vulnerabilities and generate data for safety improvements.
What you would do
- Design and execute adversarial attacks through jailbreaks, prompt injections, harmful content requests, and multi-turn manipulation scenarios
- Systematically test model responses, annotate failures, and classify vulnerabilities by type, severity, and customer impact
- Develop attack strategies that are reproducible, well-documented, and grounded in structured red teaming frameworks
- Generate datasets, reports, and findings that customers can use to remediate and prevent future failures
- Communicate vulnerability descriptions and business risk implications to both technical and non-technical leadership
Who they want
- Native fluency in English and Finnish required for this role
- Demonstrated experience with red teaming, adversarial testing, cybersecurity, or related security assessment work
- Adversarial mindset coupled with structured approach; ability to balance creative attacks with systematic methodology
- Strong communication skills to document findings clearly and translate technical details for different audiences
- Comfort engaging with sensitive content topics within ethical guidelines and company safeguards
Main skills
What the interview asks about
1.Multi-step conversation attack design
Real attacks often unfold over multiple turns; testing this reveals whether model safety measures degrade under sustained pressure.
For example: “Design a four-turn conversation sequence where each turn builds on the previous response, gradually steering the model toward harmful content. What signals indicate success?”
2.Language-specific vulnerability hunting
Models trained primarily on English may have inconsistent safety behaviors across Finnish; finding these gaps directly improves customer trust.
For example: “You suspect the model handles safety differently in Finnish versus English requests. What testing strategy would isolate this behavior, and what findings would indicate a real problem?”
3.Severity assessment and prioritization
Not all vulnerabilities require immediate fixes; distinguishing high-risk from low-risk findings focuses customer remediation on impact.
For example: “You've found that the model generates factually incorrect advice about mental health in both languages, but only harmful content generation occurs in Finnish. How do you prioritize these for your customer?”
4.Reproducible attack documentation
If attacks can't be reproduced by the customer's engineering team, they can't be fixed; clarity of documentation directly enables remediation.
For example: “You discovered a failure mode through a 15-turn conversation. Document it in a way that a customer engineer could reproduce it in 10 minutes, without ambiguity.”
5.Stakeholder-specific reporting
Security reports must resonate with executives (risk/cost) and engineers (technical details); the candidate must tailor findings appropriately.
For example: “You've discovered 30 unique vulnerabilities. Write a one-page executive summary that prioritizes what leadership needs to know, then write detailed technical findings for the engineering team.”
A task you may get
Design a three-to-five turn attack targeting the model's ability to distinguish harmful requests in Finnish versus English, document your approach and success criteria, then estimate the business impact of this vulnerability for a customer.
How to prepare
- Study how conversational AI systems implement safety measures and where those measures may fail under sustained or sophisticated attacks
- Review one comprehensive red teaming report from a published source to understand documentation standards and risk classification methods
- Practice writing technical findings for two audiences simultaneously: engineers who need reproducibility and executives who need risk narratives
The facts
- Pay
- $48–62/hr
- Hours
- Hourly, 40 hours a week
- Where
- Remote
- Field
- Miscellaneous
- Project name
- Neon
- Posted
- 7/30/2026
- Places left
- 20
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.