Training Turk

Forensic Scientist — Nuclear Materials

$65–75/task · Mercor · Task based

An expert in nuclear materials who red-teams AI systems by writing and evaluating prompts on sensitive technical topics.

What you would do

  • Write challenging prompts at three risk levels to test how AI systems handle sensitive nuclear materials questions
  • Evaluate AI responses against a defined policy standard to judge whether they correctly balance openness with safety
  • Compose reference answers explaining the technical reasoning and policy rationale for correct responses
  • Draw and defend distinctions between routine professional questions and those seeking dangerous information
  • Document your judgments in clear writing for non-specialist reviewers

Who they want

  • Deep hands-on experience characterizing real nuclear materials through detection, attribution, or analytical work
  • Background in nuclear forensics casework, radiochemistry, mass spectrometry, detection technology, or post-detonation research
  • Strong technical writing ability and prior experience with published research, technical reports, or expert testimony
  • Ability to sustain focus on misuse scenarios and the judgment to recognize when information crosses into dangerous territory
  • Experience with red-teaming, threat modeling, or security evaluation in sensitive technical domains

Main skills

Nuclear materials forensicsRed team AI evaluationDual use technology assessment

What the interview asks about

  1. 1.Distinguishing legitimate from dangerous questions

    Your core responsibility is drawing lines others cannot. You must instantly recognize when someone is using standard questions as cover for weapon-development guidance.

    For example: “A prompt asks about isotope enrichment monitoring at 3% vs 20%. The model gives separation factors and facility design. Is this routine guidance or detection-evasion information warranting refusal?”

  2. 2.Response evaluation and policy interpretation

    Judgment consistency matters. You must apply your domain expertise to assess whether a model overrefused (failing as a technical tool), underrefused (creating misuse risk), or got it right.

    For example: “An AI receives a question about atmospheric monitoring of undeclared reprocessing facilities. The model refuses entirely. Is this an overrefusal that harms technical discussion?”

  3. 3.Reference answer construction

    Your written justification becomes the standard against which other experts and non-specialists judge your judgment. It must be technically sound and policy-transparent.

    For example: “A model explains Plutonium-241 age-dating but declines to discuss weapon versus reactor signatures. Was this line-drawing correct and why?”

  4. 4.Material signature and detection expertise

    You distinguish signature analysis (how we know what happened) from signature suppression (how to hide it). This distinction determines whether information is routine forensics or operationally sensitive.

    For example: “Evaluate a response about post-detonation sampling protocols: the AI provides accurate timing and contamination-avoidance guidance. How do you assess whether this balances technical accuracy with avoiding operationally sensitive collection methodology?”

  5. 5.Adversarial prompt crafting

    Generating dual-use and adversarial prompts requires you to think like a threat actor while staying within ethical bounds. Poor prompts won't test the model; prompts that cross lines create actual misuse risk.

    For example: “Craft an adversarial prompt about material interdiction testing forensic attribution without revealing signature masking. How would you probe without teaching suppression?”

A task you may get

Evaluate three AI responses to nuclear materials questions: routine, dual-use, and adversarial. For each, assess correctness and write policy-based judgments.

How to prepare

  • Review IAEA safeguards verification documentation to understand the baseline for routine versus sensitive information
  • Practice writing technical explanations that translate your expertise into language for policy-makers and non-scientists
  • Study examples of successful and failed red-teaming evaluations in other sensitive domains to understand line-drawing patterns
  • Prepare 2-3 writing samples showing your ability to communicate complex technical judgment clearly and concisely

The facts

Pay
$65–75/task
Hours
Task based
Where
Remote
Field
Life, Physical, and Social Science
Project name
Neon
Posted
9/14/2026
Places left
100

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.