$48–52/hr · Mercor · Part time, 7 hours a week
Bilingual Finnish evaluator assessing AI model responses to sensitive prompts and adversarial inputs for safety.
What you would do
- Write expert-level Finnish prompts addressing sensitive and complex topics
- Apply classification frameworks to categorize prompts, responses, and risk levels
- Identify adversarial techniques, attack patterns, and escalation attempts
- Document your reasoning and assessments in structured evaluation formats
- Flag content for escalation based on severity and policy violations
Who they want
- Native or near-native Finnish fluency with business-level English writing
- Bachelor's degree (completed or in progress)
- Exceptional attention to detail and strong written reasoning skills
- Mature judgment around sensitive information and potential harms
- Optional background in content moderation, trust and safety, or red-teaming
Main skills
What the interview asks about
1.Finnish language and cultural nuance
AI safety depends on recognizing cultural meanings and regional sensitivities that English speakers might miss; this protects model deployment in Finnish contexts.
For example: “A prompt in Finnish uses culturally specific terminology that could be harmless colloquial speech or a coded reference to something problematic. How would you determine which?”
2.Recognizing adversarial techniques
Users actively try to bypass safety guidelines; identifying pattern-based attacks helps improve model robustness before real misuse.
For example: “You notice three prompts that seem innocent on the surface but share a common escalation pattern. How would you document this pattern and what information would you include?”
3.Structured reasoning under ambiguity
Safety decisions often involve judgment calls; clear, consistent reasoning ensures reliable evaluation and helps train AI models on subtle distinctions.
For example: “A prompt could reasonably be interpreted two different ways. One interpretation is clearly harmless, the other touches on sensitive territory. How would you classify this and justify it?”
4.Dual-use and harm assessment
Some information has legitimate uses but also enables harm; evaluators must distinguish context-dependent risks from universally problematic content.
For example: “You're evaluating a prompt requesting technical information. Explain how you'd determine whether to flag this as potentially harmful or approve it as legitimate.”
A task you may get
Evaluate 10-15 provided Finnish prompts and responses using a classification framework. For each, categorize the risk level, identify any adversarial techniques, and write clear reasoning for your judgment.
How to prepare
- Research common adversarial attack patterns used against language models and content safety systems
- Study Finnish cultural context and sensitive topics relevant to AI deployment in Finland
- Practice writing clear, structured reasoning for ambiguous judgment calls
- Review examples of strong content moderation or red-teaming documentation from similar projects
The facts
- Pay
- $48–52/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — Western Europe preferred
- Field
- Miscellaneous
- Posted
- 9/4/2026
- Places left
- 1
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.