Training Turk

Bilingual Japanese STEM Expert (PhD) — AI Safety

$68–72/hr · Mercor · Part time, 7 hours a week

A PhD scientist with native Japanese capability evaluates how AI models handle specialized chemistry and biology topics.

What you would do

  • Craft technical prompts in Japanese spanning chemistry, biology, and nuclear domains to systematically probe model capabilities.
  • Evaluate and score model responses for scientific soundness, sensitivity handling, and appropriateness of dual-use information.
  • Apply structured assessment frameworks to classify scientific conversations and content according to guidelines.
  • Leverage specialized research knowledge to judge whether outputs reflect accurate domain information.

Who they want

  • Doctoral degree in chemistry, biology, or closely related field, whether current or completed.
  • Native or near-native Japanese proficiency combined with business-quality English writing.
  • Deep working knowledge of laboratory procedures and computational methods in your research specialization.
  • Mature discernment about scientific safety, biosafety protocols, and responsible treatment of hazardous material topics.

What the interview asks about

  1. 1.Japanese scientific prompt crafting

    Creating rigorous Japanese prompts that systematically test models on technical topics is the role's core deliverable; interviewers verify depth across domains.

    For example: “If you needed to rigorously test a model's understanding of radiochemistry, what specific Japanese prompts would you construct to probe from foundational concepts through advanced edge cases involving material hazards?”

  2. 2.Dual-use risk assessment

    Evaluators must consistently distinguish legitimate scientific inquiry from information posing misuse risks; this judgment directly drives hiring decisions about model safety.

    For example: “You receive a model response in Japanese describing synthetic procedures for a controlled precursor. How would you determine whether the detail level presents concerning implications, and what assessment framework would guide your decision?”

  3. 3.Technical domain expertise

    Confidently judging model accuracy on specialized topics depends on genuine research mastery; interviewers probe the depth of your background.

    For example: “A model claims compound X undergoes hydrolysis under specific conditions yielding product Z. How would you verify this claim's validity and what laboratory context informs your evaluation?”

  4. 4.Bilingual communication precision

    Creating and evaluating scientific Japanese requires both language fluency and domain accuracy; lapses in either dimension undermine the role's impact.

    For example: “Compare how you'd explain molecular orbital theory to a general audience using Japanese versus English. What terminology obstacles emerge, and how do you maintain precision across both languages?”

  5. 5.Structured guideline application

    Consistent evaluation requires careful adherence to systematic frameworks; interviewers assess whether you follow criteria rigorously and identify ambiguous instances.

    For example: “You receive a classification schema with five sensitivity levels. A borderline prompt fits both level 2 and level 4 depending on interpretation. How would you reach a defensible determination and document your reasoning?”

A task you may get

Given a collection of Japanese chemistry prompts and corresponding model responses, classify each response on a provided sensitivity scale, identify accuracy issues, and compose an assessment of whether the model handles dual-use knowledge responsibly.

How to prepare

  • Study your most specialized research domain in depth and prepare to critique an AI model's claims in that specific area.
  • Review common dual-use considerations in chemistry and biology; develop a mental schema for distinguishing problematic versus appropriate scientific discussion.
  • Practice writing technical explanations in Japanese and English on your dissertation or research work to strengthen bilingual precision.
  • Explore how systematic rubrics function in content evaluation and safety frameworks to understand structured assessment methods.

The facts

Pay
$68–72/hr
Hours
Part time, 7 hours a week
Where
Remote · Remote — East Asia preferred
Posted
9/4/2026

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.