Training Turk

Bilingual Russian Generalist Expert — AI Safety

$38–42/hr · Mercor · Part time, 7 hours a week

A bilingual Russian-English speaker who writes safety prompts and evaluates AI model responses for harmful content and adversarial vulnerabilities.

What you would do

  • Write expert-level prompts in Russian across sensitive subject areas and dual-use topics
  • Classify model conversations and responses using structured safety guidelines
  • Flag adversarial rewording, jailbreak patterns, and escalation attempts in Russian queries
  • Document detailed reasoning for every classification and safety determination
  • Identify model weaknesses and gaps in how AI handles sensitive Russian-language content

Who they want

  • Native or near-native Russian fluency and business-level written English
  • Bachelor's degree completed or in progress
  • Strong written reasoning skills and high attention to detail
  • Sound judgment about sensitive and dual-use information handling
  • Experience with content review, red-teaming, trust and safety, or adversarial testing preferred

Main skills

Russian language fluencySafety prompt writingAdversarial pattern recognition

What the interview asks about

  1. 1.Native-level Russian prompt writing in sensitive domains

    Evaluators check whether you can write prompts that sound natural, expert, and culturally appropriate in Russian while testing AI vulnerabilities on sensitive topics.

    For example: “Write a Russian-language prompt that would naturally elicit information about dual-use chemical processes. How would you ensure the prompt sounds like something a real person would ask?”

  2. 2.Adversarial pattern recognition and rewording

    Interviewers verify you can identify how Russian speakers attempt to manipulate models through rephrasing, euphemism, or indirect framing common in Russian communication.

    For example: “An AI model refused a direct Russian request about sensitive information. How would you rephrase the request in Russian to test whether the model's safety measures are robust or triggered by keywords?”

  3. 3.Guideline application consistency and edge cases

    Evaluators test whether you apply structured guidelines uniformly and can articulate when content sits in gray areas where guidelines are unclear or conflicting.

    For example: “Your guidelines say 'flag requests for detailed information about [sensitive process].' A user asks for general background rather than detailed instructions. Would you flag it?”

  4. 4.Sensitive topic judgment and cultural context

    Interviewers assess whether you understand which topics are sensitive in Russian cultural and regulatory context, not just in English-language safety guidelines.

    For example: “A prompt discusses a technically dual-use technology but frames it in a way that is common and non-harmful in Russian academic discussion. Would you flag it as adversarial?”

  5. 5.Written classification reasoning and explainability

    Evaluators check whether your explanations are clear enough that non-Russian-speaking researchers can understand your judgment and why you classified a response a particular way.

    For example: “You flagged a model response as problematic. Your written explanation should help the AI team understand not just what was wrong, but why it matters and whether it reflects a model weakness or user manipulation.”

A task you may get

Write three Russian prompts on sensitive topics that test different safety vulnerabilities, then classify sample model responses using provided guidelines and document your reasoning for each classification.

How to prepare

  • Review the specific safety guidelines and sensitive topics you will be classifying before starting
  • Practice writing naturally in Russian on topics that are technically sensitive but commonly discussed in academic or professional contexts
  • Study how Russian communication styles differ from English in framing sensitive requests (indirectness, euphemism, context-dependence)
  • Document examples of adversarial rephrasing and jailbreak attempts you have observed in any language to understand attack patterns

The facts

Pay
$38–42/hr
Hours
Part time, 7 hours a week
Where
Remote · Remote — Eastern Europe preferred
Posted
9/4/2026

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.