$38–42/hr · Mercor · Part time, 7 hours a week
An AI safety specialist and Ukrainian language expert who strengthens models by writing test prompts and documenting how they handle sensitive topics.
What you would do
- Write expert-level Ukrainian prompts designed to test model behavior and identify vulnerabilities
- Apply guidelines to classify prompts, responses, and risk levels across different scenarios
- Document adversarial patterns, escalation sequences, and problematic model outputs
- Provide reasoning for risk judgments and communicate findings to the team
Who they want
- Native or near-native Ukrainian speaker with strong business-level English writing ability
- Bachelor's degree completed or in progress
- Excellent reasoning skills and meticulous attention to detail in evaluations
- Sound judgment on sensitive and potentially dual-use information
- Preferred: Based in Ukraine or Eastern Europe; experience with content review or red-teaming
Main skills
What the interview asks about
1.Writing prompts for adversarial testing
Safety evaluation requires crafting Ukrainian prompts that expose model vulnerabilities strategically. Interviewers assess whether candidates can design sophisticated tests rather than obvious attacks.
For example: “Design a Ukrainian prompt testing whether a model provides medical misinformation confidently. Explain your phrasing choices and why this approach would reveal the model's actual reliability versus its avoidance of clear medical claims.”
2.Classifying content and risk levels
Distinguishing harmless content from genuinely dangerous material is central to the role. This reveals risk assessment judgment and decision-making under ambiguity.
For example: “You evaluate three Ukrainian prompts about election processes. One is educational, one is advocacy, and one seeks manipulation tactics. How would you classify each, and what evidence would justify your risk ratings?”
3.Documenting ambiguous cases
Safety work often encounters edge cases without clear answers. Candidates must document ambiguity clearly so peers understand the rationale and can challenge or refine the judgment.
For example: “A model's Ukrainian response about a sensitive geopolitical topic reads as potentially educational yet contains subtle bias. Write how you'd document this case so reviewers grasp why you flagged it.”
4.Interpreting Ukrainian language context
Language and culture directly shape model outputs and safety risks. Domain expertise in Ukrainian idioms, regional usage, and cultural implications is essential to the role.
For example: “Give an example of a Ukrainian phrase that carries different risk levels depending on region or context. Describe how you'd evaluate a model's response to this ambiguity and explain what your assessment reveals.”
5.Recognizing patterns in model failures
Safety evaluation requires noticing escalation trends and repeated vulnerabilities across test cases. This skill separates thorough evaluators from those who assess cases in isolation.
For example: “After ten Ukrainian prompts on the same sensitive topic, you notice the model becomes progressively less cautious with each question. How would you document this pattern and recommend next steps?”
A task you may get
Write three Ukrainian prompts testing a sensitive topic: one baseline, one adversarial, one edge case. Classify each for risk, flag specific model outputs you'd expect to investigate, and explain your reasoning.
How to prepare
- Familiarize yourself with AI safety concepts like prompt injection, jailbreaking, and red-teaming
- Research common AI failures on sensitive topics and practice documenting risks with concrete evidence
- Prepare examples showing how Ukrainian language context and cultural knowledge affect interpretation of potentially unsafe content
- Practice explaining risk judgments concisely, using specific model outputs as evidence
The facts
- Pay
- $38–42/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — Eastern Europe preferred
- Field
- Miscellaneous
- Posted
- 9/4/2026
- Places left
- 3
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.