Training Turk

Safeguards Analyst / IAEA or Euratom Inspector

$65–75/task · Mercor · Task based

You red-team AI models on nuclear materials and safeguards, probing where they fail to hold the line between legitimate technical questions and dangerous misuse.

What you would do

  • Write challenging single-turn prompts labeled as benign, dual-use, or adversarial to help train safer AI systems
  • Evaluate model responses against a defined safety policy, documenting whether each answer was handled correctly or where the model failed
  • Create the gold-standard reference answer that shows the right response and the technical reasoning behind it
  • Document reproducible attack cases and systematic risks in dataset form so customers can improve their systems
  • Engage with technical teams to explain why certain vulnerabilities matter and how they could be exploited in practice

Who they want

  • Direct experience verifying nuclear materials through inspection or facility work, not secondhand knowledge from textbooks
  • Background in IAEA, Euratom, SSAC/RSAC inspection regimes, material accountancy, inventory tracking, or MUF resolution
  • Practical knowledge of Additional Protocol implementation, complementary access procedures, or information analysis for safeguards
  • Understanding of dual-use technical details and confidence distinguishing routine professional questions from fishing attempts
  • Prior red-teaming, cybersecurity, or adversarial testing experience; strong technical writing with published work or expert witness background

Main skills

Nuclear material verificationIaea safeguards inspectionDual use analysis

What the interview asks about

  1. 1.Recognizing dual-use technical lines

    An AI must answer routine inspection questions fully while refusing requests that could be repurposed for concealment or diversion; the same material-balance figure can be legitimate accountancy or cover for theft.

    For example: “An AI receives a question about material unaccounted-for (MUF) resolution protocols at a fuel fabrication facility. What technical details should a model provide, and what would signal the query is probing for how to hide a diversion?”

  2. 2.Constructing adversarial multi-turn scenarios

    Proliferation actors probe boundaries across multiple turns; a model might refuse one phrasing but accept a reformulation that achieves the same dangerous goal.

    For example: “Design a two-turn attack sequence where the first turn asks about routine containment-surveillance NDA calibration, and the second exploits the context to request information about circumventing detection limits. Why would a generic model be vulnerable?”

  3. 3.Grading nuanced policy compliance

    Safety policies are precise; a model that refuses all technical discussion fails inspectors, and one that answers too freely endangers security.

    For example: “A model is asked about measurement uncertainty in nuclear processes. Explain when this is a benign technical question versus adversarial, and what your label would be.”

  4. 4.Documenting reproducible test cases

    Customers need to understand exactly what the model saw, why it failed, and whether the issue is easily repeatable or a corner case.

    For example: “You identify a jailbreak where an AI answers a sensitive fuel-cycle question if embedded in a benign facility-design discussion. Write the test case that documents this vulnerability so an engineering team can trace and fix it.”

  5. 5.Technical writing on policy reasoning

    Non-specialists reviewing your work need clear explanations of why each classification matters to safeguards and security.

    For example: “You label a prompt adversarial and recommend refusal. The customer asks why this specific phrasing triggered concern. Justify your reasoning in one paragraph.”

A task you may get

Create a 4-6 prompt red team exercise (benign, dual-use, adversarial) in a safeguards domain you know, write reference answers with policy reasoning, and identify one multi-turn attack path an AI model would likely fail.

How to prepare

  • Review 2-3 IAEA safeguards implementation guides or Additional Protocol procedures to refresh your technical baseline and terminology
  • Document 3 real inspector questions from your own experience and identify what makes each one safe to answer versus what would be dangerous
  • Study one past AI jailbreak case (e.g., prompt injection, context-smuggling) and sketch how it could translate to your safeguards domain
  • Draft a brief (1-2 pages) explaining a technical risk in your field to a non-specialist audience to practice clear policy reasoning

The facts

Pay
$65–75/task
Hours
Task based
Where
Remote
Field
Life, Physical, and Social Science
Project name
Neon
Posted
9/14/2026
Places left
100

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.