$48–52/hr · Mercor · Part time, 7 hours a week
A bilingual Korean-English evaluator identifying and strengthening AI safety by testing model responses across sensitive topics.
What you would do
- Generate targeted prompts in Korean to probe AI system behavior on complex and sensitive subject matter
- Systematically categorize prompts and conversations using standardized safety evaluation frameworks
- Identify manipulation strategies, adversarial angles, and escalation tactics in attempted jailbreaks
Who they want
- Native or near-native Korean fluency with professional-grade written English capability
- Bachelor's degree completed or in active pursuit
- Strong analytical reasoning with meticulous attention to documentation and detail
- Mature judgment regarding sensitive topics, dual-use information, and content safety
- Optional: Prior experience in content review, adversarial testing, or trust and safety functions
Main skills
What the interview asks about
1.Sensitive topic navigation
Your ability to generate credible Korean-language prompts on difficult subjects reveals judgment and cultural understanding necessary for meaningful safety evaluation.
For example: “Generate three distinct Korean prompts attempting to elicit biased medical advice from an AI system, explaining why each targets a different vulnerability.”
2.Manipulation tactic analysis
Recognizing how adversaries progressively escalate requests or disguise harmful intent demonstrates critical judgment for protecting AI systems.
For example: “Here's a conversation where a user started with innocent requests then gradually shifted toward generating harmful code. Identify the escalation points and explain the progression.”
3.Language-specific evaluation
Demonstrating how Korean cultural context, phrasing conventions, or translation ambiguities affect AI safety evaluation establishes your distinctive value.
For example: “How would you evaluate whether an AI response respects cultural norms in Korean society differently than it would for English-speaking contexts?”
4.Guideline consistency application
Reliably applying evaluation frameworks and documenting reasoning ensures systematic work quality and builds confidence in your judgment.
For example: “Given these five AI model outputs on a political topic, categorize each using our framework and justify why two borderline cases fall into different risk categories.”
A task you may get
Generate five Korean-language prompts targeting different AI vulnerabilities, classify them by risk category, and document the safety reasoning for each classification.
How to prepare
- Review documentation on AI safety concepts, bias types, and dual-use information categories without requiring technical depth
- Practice identifying manipulation patterns and escalation sequences in existing conversations or chat logs
- Consider how Korean language nuances, cultural values, and translation issues differ from English evaluation
- Prepare examples from your experience with written content review, moderation, or evaluating information quality
The facts
- Pay
- $48–52/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — East Asia preferred
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.