Training Turk

Bilingual Portuguese STEM Expert (PhD) — AI Safety

$50–54/hr · Mercor · Part time, 7 hours a week

You evaluate and strengthen AI model safety in Portuguese by writing expert scientific prompts, assessing accuracy and dual-use risk, and applying structured safety guidelines.

What you would do

  • Write sophisticated test prompts in Portuguese across chemistry, biology, and radiological domains that explore model behavior on specialized technical topics
  • Evaluate model responses for scientific accuracy, alignment with current laboratory practice, and responsible handling of sensitive information
  • Apply detailed safety and accuracy guidelines to classify, flag, and document concerning outputs
  • Assess dual-use potential and recognize when scientific information requires heightened caution
  • Document evaluations with clear reasoning tied to scientific principles and safety frameworks

Who they want

  • PhD in chemistry, biology, or closely related field, completed or in progress
  • Native or near-native fluency in Portuguese with business-level written English proficiency
  • Deep familiarity with modern laboratory techniques and computational methods in your scientific subfield
  • Demonstrated depth in a specific subfield like organic or inorganic synthesis, molecular biology, microbiology, or virology
  • Sound judgment about scientific safety, dual-use risks, and responsible information handling

Main skills

Phd level scienceChemistry expertiseBiology expertise

What the interview asks about

  1. 1.Expert prompt design in Portuguese

    Test prompts must probe genuine model capabilities at expert level; superficial or incorrectly framed questions fail to reveal safety weaknesses in specialized domains.

    For example: “Write a Portuguese prompt about chemical synthesis that a genuine researcher might pose, designed to see if the model appropriately qualifies safety considerations or omits them. Explain what makes this prompt reveal meaningful model behavior.”

  2. 2.Scientific accuracy and current practice

    Models must reflect actual contemporary laboratory methods and safety conventions; outdated or theoretically incorrect information trains the wrong standards.

    For example: “A model provides Portuguese-language guidance on a synthesis procedure you know has been superseded by safer methods in the past five years. How would you evaluate this response, and what would you flag?”

  3. 3.Dual-use recognition and judgment

    Scientific information exists on a spectrum from general to dual-use; sound judgment distinguishes information that's sensitive but appropriate to discuss from content creating genuine risk.

    For example: “A model responds to a question about purification techniques that have legitimate research applications but could also enable weaponization of a dangerous substance. How would you evaluate the response, and what specifically would concern you?”

  4. 4.Subfield depth in assessment

    Your specialized knowledge in your subfield surfaces subtle errors, outdated assumptions, and safety oversights that a generalist would miss; this depth is core to your value.

    For example: “A model response about your specific subfield contains technically accurate basic information but misses a critical safety constraint or regulatory requirement you know from practice. How would you frame this feedback?”

  5. 5.Documentation of scientific reasoning

    Your written assessments become part of model training data; clear articulation of scientific and safety reasoning ensures researchers can act on your judgments.

    For example: “Write a one-paragraph assessment of a model response to a Portuguese chemistry question, explaining both scientific and dual-use considerations and why the model's answer is or isn't appropriate.”

A task you may get

You receive 3-4 Portuguese prompts on chemistry, biology, or radiological topics and produce structured evaluations assessing scientific accuracy, identifying potential dual-use concerns, applying safety guidelines, and documenting reasoning for all judgments.

How to prepare

  • Deepen your expertise in your specific subfield: review recent literature, best-practice guides, and safety publications to maintain current knowledge
  • Study published work on dual-use research governance and recognized frameworks for assessing when scientific information requires enhanced caution
  • Review safety evaluation guidelines and calibration examples from the project so you understand how your scientific judgment should map to structured assessment categories
  • Prepare to explain your research background and the specific aspects of your subfield where your expertise will strengthen model evaluation

The facts

Pay
$50–54/hr
Hours
Part time, 7 hours a week
Where
Remote · Remote — Western Europe preferred
Posted
9/4/2026

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.