Training Turk

Bilingual Ukrainian Generalist Expert — AI Safety

$38–42/hr · Mercor · Part time, 7 hours a week

An AI safety specialist and Ukrainian language expert who strengthens models by writing test prompts and documenting how they handle sensitive topics.

What you would do

  • Write expert-level Ukrainian prompts designed to test model behavior and identify vulnerabilities
  • Apply guidelines to classify prompts, responses, and risk levels across different scenarios
  • Document adversarial patterns, escalation sequences, and problematic model outputs
  • Provide reasoning for risk judgments and communicate findings to the team

Who they want

  • Native or near-native Ukrainian speaker with strong business-level English writing ability
  • Bachelor's degree completed or in progress
  • Excellent reasoning skills and meticulous attention to detail in evaluations
  • Sound judgment on sensitive and potentially dual-use information
  • Preferred: Based in Ukraine or Eastern Europe; experience with content review or red-teaming

Main skills

Bilingual ukrainian englishAI safety evaluationPrompt engineering

What the interview asks about

  1. 1.Writing prompts for adversarial testing

    Safety evaluation requires crafting Ukrainian prompts that expose model vulnerabilities strategically. Interviewers assess whether candidates can design sophisticated tests rather than obvious attacks.

    For example: “Design a Ukrainian prompt testing whether a model provides medical misinformation confidently. Explain your phrasing choices and why this approach would reveal the model's actual reliability versus its avoidance of clear medical claims.”

  2. 2.Classifying content and risk levels

    Distinguishing harmless content from genuinely dangerous material is central to the role. This reveals risk assessment judgment and decision-making under ambiguity.

    For example: “You evaluate three Ukrainian prompts about election processes. One is educational, one is advocacy, and one seeks manipulation tactics. How would you classify each, and what evidence would justify your risk ratings?”

  3. 3.Documenting ambiguous cases

    Safety work often encounters edge cases without clear answers. Candidates must document ambiguity clearly so peers understand the rationale and can challenge or refine the judgment.

    For example: “A model's Ukrainian response about a sensitive geopolitical topic reads as potentially educational yet contains subtle bias. Write how you'd document this case so reviewers grasp why you flagged it.”

  4. 4.Interpreting Ukrainian language context

    Language and culture directly shape model outputs and safety risks. Domain expertise in Ukrainian idioms, regional usage, and cultural implications is essential to the role.

    For example: “Give an example of a Ukrainian phrase that carries different risk levels depending on region or context. Describe how you'd evaluate a model's response to this ambiguity and explain what your assessment reveals.”

  5. 5.Recognizing patterns in model failures

    Safety evaluation requires noticing escalation trends and repeated vulnerabilities across test cases. This skill separates thorough evaluators from those who assess cases in isolation.

    For example: “After ten Ukrainian prompts on the same sensitive topic, you notice the model becomes progressively less cautious with each question. How would you document this pattern and recommend next steps?”

A task you may get

Write three Ukrainian prompts testing a sensitive topic: one baseline, one adversarial, one edge case. Classify each for risk, flag specific model outputs you'd expect to investigate, and explain your reasoning.

How to prepare

  • Familiarize yourself with AI safety concepts like prompt injection, jailbreaking, and red-teaming
  • Research common AI failures on sensitive topics and practice documenting risks with concrete evidence
  • Prepare examples showing how Ukrainian language context and cultural knowledge affect interpretation of potentially unsafe content
  • Practice explaining risk judgments concisely, using specific model outputs as evidence

The facts

Pay
$38–42/hr
Hours
Part time, 7 hours a week
Where
Remote · Remote — Eastern Europe preferred
Field
Miscellaneous
Posted
9/4/2026
Places left
3

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.