$48–52/hr · Mercor · Part time, 7 hours a week
You evaluate how AI systems handle sensitive Japanese-language prompts and conversations to strengthen their safety.
What you would do
- Compose challenging prompts in Japanese targeting various sensitive domains and edge cases
- Classify AI model responses against a structured rulebook with assigned error codes
- Identify and flag language patterns that suggest attempts to manipulate or adversarially test the system
- Document justifications for each judgment in clear language
- Tag escalation patterns where conversations shift toward harmful content
Who they want
- Bilingual fluency in Japanese (native or near-native) and professional English
- Bachelor's degree completed or in progress
- Careful attention to detail and sound judgment on sensitive subjects
- No prior AI or machine learning background required
What the interview asks about
1.Recognizing indirect harmful requests
Adversarial users often phrase sensitive requests indirectly; catching these requires cultural knowledge and pattern recognition specific to this role.
For example: “An AI responded to a Japanese prompt about medicine synthesis. The request seemed academic on the surface. How would you decide if it crossed into unsafe territory?”
2.Applying error taxonomy consistently
Reviewers must tag responses using a fixed classification system hundreds of times; inconsistency degrades training data quality.
For example: “Two similar model outputs fail your safety check. Do you mark both with the same error code or different ones, and how do you justify the choice?”
3.Translating cultural judgment into standards-based assessment
Safety guidelines are often Western-centric; this role bridges cultural context to universal evaluation criteria.
For example: “A Japanese response uses culturally-specific indirect phrasing that isn't inherently unsafe. How do you score it against a rulebook written in English?”
4.Detecting escalation sequences in multi-turn conversations
Single-turn toxicity detection misses campaigns where users gradually push boundaries across multiple messages.
For example: “In a 5-message exchange, the first two seem harmless in Japanese but the final three requests shift focus. How do you document the escalation pattern?”
5.Balancing false positives and false negatives
Rejecting safe content wastes training resources; missing genuinely unsafe content undermines safety goals; this judgment shapes model behavior.
For example: “A response about a sensitive pharmaceutical use case is technically safe but easily misinterpreted. Do you flag it as ambiguous or pass it through?”
A task you may get
Grade 5-10 sample prompts and responses using a provided rubric, then peer-review one other contractor's scorecards and flag any inconsistencies.
How to prepare
- Prepare examples of indirect phrasing in Japanese that can mask harmful intent
- Understand the difference between cultural sensitivity and safety risk categories
- Read through the error taxonomy and anticipate which codes apply to different scenarios
- Practice writing structured justifications that could be understood by non-Japanese speakers
The facts
- Pay
- $48–52/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — East Asia preferred
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.