$58–62/hr · Mercor · Part time, 7 hours a week
A PhD-credentialed scientist fluent in German and English who protects AI model safety by designing specialized queries and assessing responses.
What you would do
- Create expert-caliber German prompts probing model performance across chemistry, molecular pathways, radiological topics, and nuclear phenomena
- Analyze generated outputs for scientific validity and contemporary laboratory practice alignment
- Evaluate whether responses handle restricted chemical procedures, biosecurity considerations, or dual-use applications with adequate responsibility
- Apply systematic categorization to classify model behavior relative to safety benchmarks
- Prepare comprehensive assessments detailing your findings and disciplinary rationale
Who they want
- PhD graduate in chemistry, biology, radiochemistry, physics, or closely allied STEM fields, either completed or currently pursuing
- Native or fluent-level German communication with capable business English proficiency across spoken and written contexts
- Hands-on research experience or computational modeling background demonstrating practical facility with methods and equipment in your area
- Professional judgment regarding scientific safety protocols, biosafety regulations, and responsible stewardship of sensitive technical material
- Part-time contract availability at 58-62 USD hourly with immediate or flexible start date feasible
Main skills
What the interview asks about
1.Constructing rigorous German chemistry prompts
Interviewers test your ability to generate German-language questions that authentically challenge model comprehension in specialized domains, balancing disciplinary precision with linguistic fluency.
For example: “Design a German prompt examining synthesis procedures for a complex molecule that would realistically stress-test model understanding of safety protocols. What German terminology choices make this prompt particularly effective for evaluation?”
2.Assessing research methodology accuracy
Your authority derives from recognizing when models misrepresent techniques, misapply procedures, or cite outdated protocols, requiring intimate familiarity with your field's evolution.
For example: “A response describes a radiochemical procedure using outdated safety practices from the 1980s instead of modern protocols. How would you evaluate this response, and what written documentation would you provide?”
3.Identifying inadequate safeguarding in responses
This position centers on recognizing unsafe information disclosure where models fail to appropriately restrict dangerous knowledge, requiring research judgment and discretion.
For example: “A model response covers biosecurity-sensitive topics with excessive mechanistic detail about production pathways. Describe your assessment approach: what triggers concern versus what remains acceptable scientific discussion?”
4.Operating systematic classification frameworks
Uniform classification across numerous examples produces reliable pattern recognition that supports algorithmic improvement and comparative model analysis.
For example: “You've assessed 12 radiological topic exchanges with inconsistent model performance on safety considerations. How would you organize your findings to isolate whether the model has knowledge deficiency versus judgment problems?”
5.Managing language precision across translation
Switching between German and English creates susceptibility to meaning shifts, so bilingual fluency prevents assessment errors rooted in linguistic ambiguity instead of actual model failure.
For example: “A German prompt employs specialized terminology with no exact English equivalent, and the model's response reveals confusion stemming from this linguistic gap. How would you structure your documentation to explain the language-mediated issue?”
How to prepare
- Compile a summary of your dissertation or principal research focusing on methods, instrumentation, and highest-risk materials or techniques involved
- List and briefly justify five contemporary subject areas you regard as presenting dual-use or biosecurity concerns within your discipline
- Gather laboratory safety and biosafety training records, preparing to discuss scenarios showing your responsible judgment
The facts
- Pay
- $58–62/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — Western Europe preferred
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.