$61–65/hr · Mercor · Part time, 7 hours a week
A doctoral-level STEM researcher with fluent Finnish and English evaluates AI model responses to specialized scientific prompts, assessing accuracy, safety, and training effectiveness.
What you would do
- Compose expert-level scientific prompts in Finnish across chemistry, biology, and related specialized fields
- Evaluate AI responses for technical accuracy, methodological soundness, and alignment with current research practice
- Classify conversations and prompts according to structured annotation guidelines
- Identify and escalate safety concerns related to dual-use information or sensitive technical topics
- Provide detailed assessment of model knowledge gaps and learning opportunities in your scientific subfield
Who they want
- PhD in chemistry, biology, or closely related STEM field (completed or in progress)
- Native or near-native Finnish fluency with business-level written English ability
- Deep practical knowledge of modern laboratory and computational techniques in your area
- Sound judgment about scientific safety, biosafety implications, and dual-use considerations
- Preferred: background in chemical safety, biosecurity, technical content review, or hazardous-materials handling
Main skills
What the interview asks about
1.Dual-use scientific assessment
AI must learn to balance providing legitimate scientific education against enabling harmful applications; your judgment determines whether responses appropriately navigate this tension
For example: “The model provides a detailed synthesis procedure for a common chemical that could have dual-use implications. How would you evaluate whether this response crosses from educational to dangerous, and what specific feedback would guide model improvement?”
2.Finnish technical precision
AI trained on multilingual data needs evaluators who catch terminology ambiguities in non-English languages; mistakes here affect all scientists working in Finnish
For example: “A model uses two different Finnish terms for the same procedure interchangeably. One is the modern standard, the other is outdated. What signal does this inconsistency send about the model's Finnish training, and how would you document this?”
3.Outdated versus incorrect science
Evaluators must distinguish between responses using older methods that still work versus responses with factual errors; this difference guides whether to flag the response and how
For example: “The model describes a sterilization protocol from nuclear chemistry that's no longer standard practice but remains technically valid. Would you annotate this as outdated, acceptable, or a gap in model training? How would you frame this?”
4.Guideline ambiguity resolution
Classification rules sometimes conflict with scientific reality; your expertise reveals these tensions so the team can refine guidelines rather than force-fit annotations
For example: “You're classifying a radiochemistry discussion where the safety category seems to conflict with legitimate educational content the model produced. Do you override the guideline or note the conflict? What additional context would you provide?”
A task you may get
Given 4-5 sample AI responses to Finnish prompts in your doctoral subfield covering chemical safety or biosafety topics, annotate each for scientific accuracy, safety implications, and correct classification using provided guidelines.
How to prepare
- Review how machine learning systems improve through evaluator feedback during training
- Gather examples of scientific concepts that non-experts often misunderstand or that AI frequently gets wrong
- Refresh your knowledge of current safety standards and practices in your specific STEM field
- Research how dual-use information is typically handled in AI safety evaluation frameworks
The facts
- Pay
- $61–65/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — Western Europe preferred
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.