Training Turk

AI Safety Experts — English & Punjabi

$16–22/hr · Mercor · Hourly, 40 hours a week

An AI safety specialist who probes conversational models with adversarial inputs to surface vulnerabilities and generate training data.

What you would do

  • Design and execute red team attacks on AI models targeting jailbreaks, prompt injections, misuse cases, and bias exploitation
  • Generate high-quality labeled data that captures failure modes, classifying vulnerabilities and flagging systemic risks
  • Apply published taxonomies and testing playbooks to ensure consistency and rigor across all assessment work
  • Document reproducible attack scenarios, datasets, and findings so downstream teams can strengthen models

Who they want

  • Fluent in both English and Punjabi with native or near-native command of nuance and cultural context
  • Strong judgment about language: ability to explain why an AI response succeeds or fails, with clear reasoning
  • Rigorous attention to detail: you catch inconsistencies and logical gaps that less careful reviewers miss
  • Experience communicating complex findings to both technical and non-technical stakeholders in writing

Main skills

Adversarial prompt engineeringBias detection and analysisVulnerability classification

What the interview asks about

  1. 1.Adversarial prompt design

    Your ability to construct unexpected inputs that bypass safety guardrails determines how well your testing uncovers real vulnerabilities.

    For example: “Walk me through how you would design a multi-turn conversation that gradually escalates requests to push a model toward generating harmful content, and how you'd document each turning point.”

  2. 2.Vulnerability categorization

    Distinguishing between critical failures and benign edge cases is essential so clients prioritize fixes effectively and don't waste resources on low-risk findings.

    For example: “You're testing a model trained to refuse medical advice. It declines to diagnose symptoms but suggests three over-the-counter remedies for someone describing clear pneumonia signs. How would you classify this failure and explain why it matters?”

  3. 3.Bias detection in context

    Bias often emerges through subtle patterns across many queries rather than a single failure, and your ability to surface these trends shapes how clients understand their model's blind spots.

    For example: “Across 50 job recommendation queries, the model ranks males higher for engineering roles, females higher for HR. How would you present this pattern? What follow-up testing would you recommend?”

  4. 4.Documentation for technical audiences

    Your written explanations are how engineers understand which inputs triggered failures and how to reproduce them, directly determining whether fixes actually address root causes.

    For example: “Describe how you would document a jailbreak attempt that required 8 steps to succeed, including the exact model responses at each stage, so another researcher can verify your work and an engineer can trace the vulnerability.”

  5. 5.Cross-cultural language judgment

    Your bilingual fluency lets you catch culturally contextual biases and meaning shifts that monolingual testers miss, which is essential for models serving diverse user bases.

    For example: “A model trained primarily on English data sometimes generates responses in Punjabi that are technically grammatical but culturally inappropriate or offensive. How would you systematically surface these issues in your testing report?”

A task you may get

Conduct a simulated red team assessment of a conversational AI model, crafting adversarial prompts targeting three distinct vulnerability classes, documenting each attempt with the model's response and your classification of the failure mode.

How to prepare

  • Study published AI safety benchmarks and red team frameworks to understand standard taxonomies and testing methodologies
  • Prepare examples from your experience where you identified subtle errors or misalignments in language or logic that others overlooked
  • Review case studies of real AI failures in bias, hallucination, or harmful output generation to ground your understanding of stakes

The facts

Pay
$16–22/hr
Hours
Hourly, 40 hours a week
Where
Remote
Field
Data Analysis
Project name
Neon
Posted
6/4/2026
Places left
1

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.