Training Turk

AI Safety Experts — English & Kannada

$16–22/hr · Mercor · Hourly, 40 hours a week

A bilingual safety tester who probes AI models with adversarial inputs to uncover vulnerabilities and generate data that makes systems more robust.

What you would do

  • Test conversational AI with jailbreaks, prompt injections, bias exploitation, and multi-turn manipulation tactics
  • Classify discovered vulnerabilities using structured taxonomies and quality standards
  • Generate high-fidelity human evaluation data annotating model failures and edge cases
  • Produce reproducible test cases and documented attack scenarios customers can use for remediation

Who they want

  • Native or fluent proficiency in both English and Kannada
  • Strong judgment about language, content accuracy, and appropriateness
  • Rigor and attention to detail, noticing errors others overlook
  • Comfort working to structured guidelines and quality standards consistently
  • Adaptability across different project types and customer requirements

What the interview asks about

  1. 1.Jailbreak technique effectiveness

    Understanding what prompting strategies actually bypass safety measures helps systems learn to defend against them.

    For example: “You tested three jailbreak tactics: role-playing as a harmful character, gradual escalation, and language switching. Which proved most effective and why did Kannada responses differ from English ones?”

  2. 2.Bias and fairness detection

    Models can produce fair-seeming outputs that contain subtle stereotypes or discriminatory patterns only visible to expert judgment.

    For example: “You asked the AI to describe engineers, finance candidates, and decision-makers. Responses seemed neutral but showed gender stereotyping patterns. How did you classify and document that vulnerability?”

  3. 3.Consistency across languages

    Safety measures sometimes apply inconsistently across languages, and AI models need human experts to catch these asymmetries.

    For example: “A harmful request was properly refused in English but accepted when you rephrased it in Kannada. Did the system lack training data in Kannada or was it a different architectural issue, and how did you report that finding?”

  4. 4.Multi-turn manipulation detection

    Some vulnerabilities require conversation over multiple turns, building context that single-exchange testing misses.

    For example: “You started a benign conversation about a sensitive topic, gradually shifted context over five turns, and the model eventually provided harmful content it initially refused. How did you structure that test case for customer reproduction?”

  5. 5.Documentation rigor

    Customers can only act on vulnerabilities if your reports are specific, reproducible, and properly classified for triage.

    For example: “You found 12 instances of inconsistent refusal behavior across similar prompts. How did you organize those findings so a customer could efficiently prioritize which to address first?”

A task you may get

Test a conversational AI system with three different adversarial strategies, classify the vulnerabilities you discover, and document one reproducible attack case with exact prompts and expected outcomes.

How to prepare

  • Identify three jailbreak or prompt injection techniques from public research and practice applying them to understand what makes them effective
  • Study the provided taxonomy for classifying AI safety failures so you can apply it consistently
  • Prepare examples of subtle bias or inconsistency in AI outputs from your own interactions with similar systems
  • Review documentation examples to understand the level of specificity customers need for reproducibility

The facts

Pay
$16–22/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Miscellaneous
Project name
Neon
Posted
6/4/2026
Places left
2

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.