$38–42/hr · Mercor · Part time, 7 hours a week
You write Croatian prompts and evaluate whether AI models handle sensitive topics safely, flagging escalation patterns and dual-use risks.
What you would do
- Write expert-level prompts in Croatian designed to test AI responses across sensitive subject areas, leveraging cultural knowledge and linguistic authenticity
- Apply structured classification guidelines to evaluate both prompts and model responses, determining appropriate safety categorization
- Identify adversarial phrasing patterns, social engineering tactics, and escalation strategies that attempt to elicit unsafe outputs from AI systems
- Document your reasoning for each safety judgment, explaining why a response is acceptable, needs guardrails, or demonstrates failure
- Provide detailed feedback to help AI systems learn appropriate handling of cultural nuance, context-dependent harms, and dual-use information in Croatian
Who they want
- Native or near-native Croatian fluency combined with business-level written English proficiency
- Bachelor degree completed or in progress in any discipline
- Demonstrated strong written reasoning skills and exceptional attention to detail in evaluating subtle language and context
- Sound judgment about sensitive information, dual-use risks, and when information is harmful versus legitimate depending on context
- Experience in trust and safety, content review, adversarial testing, or policy evaluation is valuable but training is provided
What the interview asks about
1.Culturally-Aware Prompt Authoring
Safety testing must account for cultural context and linguistic expression; interviewers assess whether you write authentic Croatian prompts rather than relying on literal translations.
For example: “You write adversarial prompts about online fraud targeting elderly people. A direct English translation would not work in Croatian culture. How would you write culturally appropriate prompts testing elder vulnerability understanding?”
2.Dual-Use Information Discrimination
Some information is legitimate in one context and harmful in another; interviewers evaluate whether you distinguish based on stated intent and context rather than blanket restrictions.
For example: “A request asks for information about chemical reactions that could be used legitimately for industrial safety or farming, but also could theoretically enable harm. How do you evaluate what the AI should do, and how does the Croatian context matter?”
3.Escalation Pattern Detection
Adversarial actors use step-by-step approaches to bypass guardrails; interviewers probe whether you recognize patterns where early requests seem innocent but build toward unsafe territory.
For example: “A conversation starts with factual questions about political movements, shifts to grievances and scapegoating, then hints at confrontation. At what point would you flag this escalation and explain the progression?”
4.Classification Consistency and Reasoning
Your judgments train AI models; interviewers assess whether your classifications are consistent and your reasoning would allow another Croatian speaker to reach the same conclusion.
For example: “You have classified three different model responses to sensitive prompts into 'caution required,' 'acceptable with caveats,' and 'clear refusal needed.' Explain your classification criteria so someone else could apply the same standard to new cases.”
5.Nuanced Safety Feedback Articulation
AI systems learn from explanation, not just labels; interviewers evaluate whether you articulate specifically why a response succeeds or fails and what the model should improve.
For example: “An AI response in Croatian is factually reasonable but uses language that could unintentionally normalize harm. How would you explain this subtle failure so it learns accuracy versus appropriateness distinction?”
6.Context-Dependent Information Risk Assessment
Harm depends on context and audience; interviewers probe whether you evaluate requests differently based on apparent legitimate use versus suspicious patterns.
For example: “A user requests information about psychological manipulation tactics. If they frame it as academic research about advertising ethics, does that change your safety classification compared to a request without context? How would you evaluate the difference?”
A task you may get
Write three expert-level Croatian prompts across sensitive subject areas testing AI safety. Classify model responses using guidelines. Document reasoning and flag escalation patterns. Explain how cultural context informed your prompts.
How to prepare
- Review examples of adversarial testing and red-teaming frameworks to understand how safety specialists approach prompt design and response evaluation
- Consider which sensitive topics are culturally specific to Croatian context versus universal, and how phrasing changes meaning across cultures
- Study how social engineering and escalation tactics work in conversation, noting how they might manifest differently in Croatian discourse
- Document your own experience with online safety, content moderation, or recognizing manipulation tactics, even informally from social media or community experience
The facts
- Pay
- $38–42/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — Eastern Europe preferred
- Field
- Miscellaneous
- Posted
- 9/4/2026
- Places left
- 1
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.