Training Turk

Health Physicist

$65–75/task · Mercor · Task based

A radiological safety expert writes adversarial prompts and evaluates AI model responses for misuse potential in nuclear science.

What you would do

  • Write multi-level adversarial prompts (benign, dual-use, adversarial) testing AI responses in radiological safety contexts
  • Evaluate model outputs against defined policy standards for appropriate refusal or acceptance of requests
  • Develop reference answers explaining technical reasoning for why particular responses meet or fail safety standards
  • Assess whether models consistently distinguish between routine professional inquiries and dual-use research risks
  • Provide structured feedback helping improve model judgment for safety-critical domains

Who they want

  • Direct operational experience in health physics: survey work, assay, contamination control, dose planning, or facility radiation safety
  • Institutional background at DOE facility, national laboratory, reactor site, or high-hazard facility where operational errors carry serious consequences
  • Demonstrated expertise in radiological risk assessment, shielding calculations, dose reconstruction, or activation product management
  • Technical writing proficiency evidenced by published research, expert witness testimony, or professional documentation
  • Red-teaming experience or demonstrated ability to think adversarially about dual-use technical concepts and misuse scenarios

Main skills

Radiation safety assessmentDual use research detectionPrompt engineering and adversarial testing

What the interview asks about

  1. 1.Dual-use scenario discrimination

    Recognizing when a technical question masks misuse intent requires field experience, since the same shielding calculation protects workers or conceals sources, and AI judgment failures in this domain create safety risks.

    For example: “A question asks about gamma-ray attenuation through lead at various thicknesses. This is routine, but could mask dual-use intent. How would you reframe it as an adversarial prompt?”

  2. 2.Prompt engineering at the boundary

    Effective red-teaming requires writing questions that sit exactly on the policy boundary, neither clearly safe nor obviously dangerous, forcing the model to apply judgment rather than pattern-matching.

    For example: “Write two versions of a neutron activation question: one appropriate and one probing dual-use concerns. Explain what makes the adversarial version test the model's judgment.”

  3. 3.Reference answer clarity for non-specialists

    Your written reasoning must explain radiological concepts to policy experts without physics backgrounds, so they can evaluate whether the model's reasoning was sound and whether your answer is appropriate.

    For example: “Write a two-sentence explanation for a policy reviewer why bioassay dose reconstruction methodology is a legitimate professional question, not a misuse risk.”

  4. 4.Operational context and real-world constraints

    Answers that miss practical field constraints (equipment capabilities, regulatory limits, time constraints) signal the model lacks genuine operational grounding, raising questions about whether it understands risk appropriately.

    For example: “A model recommends decontamination procedures requiring equipment not typically deployed in contaminated sites. How do you evaluate whether this is sound judgment or impractical reasoning?”

A task you may get

Write two prompts in radiological shielding or dose planning (benign and adversarial). Develop reference answers explaining why each should be handled differently. Assess what model behavior indicates sound judgment versus problems.

How to prepare

  • Review recent dual-use research policy discussions in radiological science to understand current regulatory framing and misuse concerns
  • Prepare examples from your operational experience of legitimate questions you've received that could be misframed as dual-use requests
  • Study at least two published research papers in your health physics specialty, noting how technical content could have legitimate and problematic applications
  • Reflect on boundary cases you've encountered: situations where a client's request seemed ambiguous or where you needed judgment to determine appropriateness

The facts

Pay
$65–75/task
Hours
Task based
Where
Remote
Field
Life, Physical, and Social Science
Project name
Neon
Posted
9/14/2026
Places left
100

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.