Training Turk

AI Safety Experts — English & Indonesian

$17–25/hr · Mercor · Hourly, 40 hours a week

You red-team AI models across English and Indonesian, uncovering vulnerabilities through adversarial prompts and generating datasets for safer AI deployment.

What you would do

  • Design and execute red-team attacks on conversational AI including jailbreaks, prompt injection, and adversarial reasoning exploits
  • Annotate and classify model failures using structured taxonomies, documenting vulnerability type, severity, and root cause for each case
  • Generate high-quality adversarial datasets with reproducible test cases, clear labels, and policy documentation that technical teams can use for model improvement
  • Probe culturally-grounded vulnerabilities that monolingual testing misses - bias patterns, social manipulation tactics, and regional language-specific exploits
  • Communicate findings to engineering and non-technical stakeholders, explaining why each vulnerability poses a real risk and what customer impact it creates

Who they want

  • Proven red-teaming, adversarial-ML, or cybersecurity experience where you conducted systematic vulnerability probing and documented findings
  • Curious and adversarial-minded: you think creatively about how to break systems and find unconventional attack paths
  • Structured methodology: you follow benchmarks, taxonomies, and frameworks rather than pursuing random vulnerabilities
  • Native fluency in English and Indonesian (global, excluding Brazilian Portuguese); ability to communicate technical findings clearly in both languages
  • Adaptability to work across multiple projects and customers, switch contexts rapidly, and demonstrate resilience during sustained engagement with sensitive content

Main skills

Conversational AI jailbreakingMultilingual bias detectionAdversarial prompt design

What the interview asks about

  1. 1.Bilingual attack vector discovery

    Vulnerabilities that work in one language may not translate directly; true multilingual red teamers identify language-specific exploits others miss.

    For example: “A model is trained on balanced English-Indonesian data. Design an attack that exploits English-language jailbreak techniques but adapts it to Indonesian idioms or cultural references that might be more effective in that language.”

  2. 2.Annotating nuanced vulnerability patterns

    Raw failure examples are not actionable; testers must identify the underlying weakness and tag it so engineers know what class of problem to fix.

    For example: “You collect three different prompts that lead a model to generate biased hiring advice. How would you annotate them so an engineer understands whether the issue is training data bias, prompt template injection, or conversational memory mishandling?”

  3. 3.Regional and cultural exploitation

    The same social manipulation tactic does not work uniformly across cultures; effective red teamers understand local context and adapt accordingly.

    For example: “In Indonesian culture, indirect communication and hierarchical respect norms differ from US English norms. How would you design a social engineering attack on an AI system that exploits these cultural patterns to bypass safety guardrails?”

  4. 4.Reproducible testing under ambiguity

    Models behave stochastically; a one-off failure is not actionable unless you can show when and why it happens consistently.

    For example: “You uncover an occasional model failure that seems tied to conversational history but is not reliable. How would you design an experiment to isolate the conditions that trigger the vulnerability?”

  5. 5.Explaining risk across languages and domains

    Your findings must be clear to both ML engineers and non-technical decision-makers in multiple language contexts.

    For example: “You discovered a bias in a customer-support chatbot where it gives lower-quality responses to Indonesian speakers. Write a 300-word explanation for an Indonesian stakeholder team that does not assume ML knowledge.”

A task you may get

Develop a structured red-team assessment creating six adversarial test cases spanning multiple attack vectors. Document each case and the model's response.

How to prepare

  • Study two published adversarial techniques such as prompt injection or role-play jailbreaks to understand how they work
  • Review recent academic or practitioner literature on multilingual AI bias to identify regional vulnerabilities in your domain
  • Collect and analyze three real-world examples of AI safety failures in customer-facing systems and identify what adversarial techniques would have surfaced each before deployment
  • Document two cultural or linguistic factors that affect how Indonesian and English speakers might interact differently with AI systems, and sketch vulnerability approaches for each

The facts

Pay
$17–25/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Project name
Neon
Posted
7/30/2026
Places left
20

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.