Training Turk

AI Safety Experts — English & Danish

$48–62/hr · Mercor · Hourly, 40 hours a week

You adversarially test conversational AI models to uncover safety vulnerabilities and generate training data for AI safety improvement.

What you would do

  • Design and execute red teaming attacks: jailbreaks, prompt injections, multi-turn manipulation, and bias exploitation strategies
  • Annotate model failures, classify vulnerabilities by type and severity, and flag systemic risks in AI responses
  • Apply structured frameworks and taxonomies to conduct consistent testing across different models and scenarios
  • Produce detailed, reproducible attack documentation and datasets that customers can validate and remediate
  • Communicate findings clearly to both technical teams and non-technical stakeholders about risk and impact

Who they want

  • Native fluency in English and Danish required for this role
  • Prior experience with red teaming, adversarial AI testing, cybersecurity, or socio-technical risk assessment
  • Curious, adversarial mindset with instinct to push systems to failure points systematically
  • Structured approach using frameworks, not random attacks; ability to document findings clearly
  • Comfortable discussing sensitive topics like bias, misinformation, and harmful content with clear ethical guidelines

Main skills

Prompt injection techniquesJailbreak methodologyBias exploitation testing

What the interview asks about

  1. 1.Jailbreak design and execution

    Demonstrates understanding of model constraints and ability to design multi-step attacks that surface inconsistencies in safety training.

    For example: “Design a three-turn conversation that gradually escalates from benign requests to increasingly harmful content. How would you measure when the model has failed?”

  2. 2.Bias discovery in multilingual contexts

    Tests whether the candidate recognizes that bias manifests differently across languages and cultural contexts, essential for Danish language testing.

    For example: “You're testing whether a model treats Danish versus English requests for sensitive topics differently. What attack patterns would you use to detect this disparity?”

  3. 3.Taxonomy and framework application

    Shows ability to work systematically rather than chaotically, which separates effective red teamers from those producing noisy, irreproducible results.

    For example: “Map a vulnerability you discovered to both the OWASP AI taxonomy and your project's internal taxonomy. Explain why multiple categorizations matter for remediation.”

  4. 4.Documentation for customer action

    Reveals whether findings are communicated as isolated examples or as reproducible, systematized attacks customers can actually fix.

    For example: “You discovered that the model fails when given role-play scenarios framed in Danish versus English. Write a report summarizing the vulnerability and reproducible steps.”

  5. 5.Prioritizing vulnerabilities by impact

    Tests judgment about which failures matter most and why, ensuring effort focuses on genuine risks rather than theoretical edge cases.

    For example: “You've found 20 potential jailbreaks this week. Your customer can fix 3 before launch. Which three do you prioritize and why?”

A task you may get

Design a three-part attack sequence targeting a conversational AI's handling of misinformation, then document one successful failure case with steps a customer engineer could reproduce and fix.

How to prepare

  • Review published frameworks for AI safety testing (OWASP, NIST, or similar) to understand structured red teaming methodologies
  • Study at least two research papers on adversarial attacks in conversational AI to learn sophisticated attack patterns
  • Practice writing technical findings so non-engineers can understand vulnerability severity and remediation options

The facts

Pay
$48–62/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Project name
Neon
Posted
7/30/2026
Places left
20

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.