$63–67/hr · Mercor · Part time, 7 hours a week
A bilingual PhD scientist creates Korean-language prompts and evaluates model responses to improve AI safety in technical domains.
What you would do
- Write scientifically sophisticated prompts in Korean designed to test model capabilities and safety on technical topics
- Evaluate model-generated responses for scientific accuracy, helpfulness, and responsible handling of sensitive subjects
- Classify prompts and conversations using structured safety and quality guidelines
- Flag responses that may mishandle dual-use information or contain factual errors
Who they want
- PhD in chemistry, biology, or closely related discipline with the depth to write expert prompts across your domain
- Native or near-native fluency in Korean combined with business-level English written communication
- Deep familiarity with modern laboratory techniques and computational methods in your subfield
- Sound judgment about scientific safety risks and responsible handling of sensitive technical content
Main skills
What the interview asks about
1.Technical domain expertise
You must recognize when a model gives scientifically incorrect answers; shallow knowledge causes you to miss errors that matter for safety.
For example: “Consider a molecular biology question about protein expression techniques. What three realistic mistakes would a model make on this topic, and why would someone with PhD-level training catch them?”
2.Korean scientific writing
You're creating training data in Korean for models; imprecise or unclear prompts produce poor quality training and confuse the model's learning.
For example: “Write a Korean-language prompt about organometallic synthesis that would test whether a model understands reagent selectivity. What technical terms did you choose and why?”
3.Dual-use risk recognition
The core safety value you add is flagging when model responses cross from legitimate science into potentially dangerous knowledge; this judgment separates qualified contributors.
For example: “A model answers a chemistry question with detailed synthetic procedures. How would you evaluate whether the response appropriately handles risk versus suppressing legitimate science education?”
4.Prompt design for testing
Quality prompts reveal model weaknesses; poorly designed prompts miss gaps in model knowledge or safety reasoning.
For example: “Design two Korean prompts on microbiology that target different types of model capability. Why would each one be effective at testing the model's performance?”
5.Cross-domain reasoning
Safety risks often emerge when dual-use information from one field combines with knowledge from another; you must recognize these patterns.
For example: “You have expertise in radiochemistry. How would safety evaluation change if the model were generating responses combining chemistry concepts with engineering or biological applications?”
A task you may get
Write 2-3 Korean-language scientific prompts in your specialization and provide analysis of what correct responses should cover and where model errors might occur.
How to prepare
- Review recent scientific literature in your specialty, noting areas where models commonly make errors or give incomplete answers
- Prepare examples of scientific concepts from your field that require nuanced explanation versus dangerous simplification
- Consider how your field's terminology translates into Korean and practice explaining technical concepts bilingually
- Reflect on safety-sensitive topics in your discipline and how you would guide a model to handle them responsibly
The facts
- Pay
- $63–67/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — East Asia preferred
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.