Training Turk

Nuclear Engineering & Safeguards Experts for Red Team

$65–75/task · Mercor · Task based, 40 hours a week

Nuclear domain expert who red-teams AI models by writing challenging prompts and evaluating whether responses correctly identify misuse potential.

What you would do

  • Craft single-turn prompts at three levels, from straightforward to adversarial, anchored in realistic technical scenarios
  • Judge whether each AI response correctly applies your organization's policy on handling requests with potential for harm
  • Write technical rationales explaining your assessment in plain language for non-specialist readers

Who they want

  • Deep expertise in fuel cycles, material safeguards, nonproliferation policy, or nuclear security from IAEA, national laboratory, or equivalent
  • Demonstrated technical writing ability through published research, expert testimony, or prior red-teaming work
  • Comfort discussing misuse scenarios within ethical bounds and ability to step back at any point without consequence

Main skills

Nuclear engineeringSafeguards analysisNonproliferation

What the interview asks about

  1. 1.Prompt design under policy constraints

    The core work is writing prompts that test a model's judgment at the policy boundary. Interviewers need to see you think through plausibility, technical accuracy, and escalation logic.

    For example: “A researcher asks how laser-enrichment systems compare to centrifuges in a fuel cycle. Write a dual-use version a model should refuse and explain why that boundary exists.”

  2. 2.Technical judgment and policy application

    Evaluating responses requires deep domain knowledge plus disciplined thinking about where a policy draws its line. This tests whether you articulate safeguards logic clearly in writing.

    For example: “You write a prompt about material accounting discrepancies in a facility. The AI provides a technically correct answer but omits detection methods and policy implications. Is that a pass, and why?”

  3. 3.Cross-domain communication

    Your written reasoning must convince smart non-specialists that your judgment is sound. This tests whether you explain nuclear reasoning without jargon or oversimplification.

    For example: “Explain to a policy analyst without nuclear knowledge why you marked a response as handling a nonproliferation concern correctly. What are the core technical facts they need?”

  4. 4.Expertise depth and background credibility

    The role's credibility depends on genuine depth in safeguards, security, or fuel-cycle work. Interviewers probe actual hands-on experience and how you developed it.

    For example: “Describe one material accounting or export control judgment call from your background where the technical detail determined the outcome.”

A task you may get

Write a dual-use prompt in your specialization, evaluate a provided AI response against a safety policy, and compose a 150-word technical rationale for whether the model passed or failed.

How to prepare

  • Review recent AI safety research on dual-use disclosure and misuse potential assessment
  • Gather 2-3 technical writing samples or publications you can reference
  • Prepare concrete examples from your safeguards work where judgment calls significantly mattered

The facts

Pay
$65–75/task
Hours
Task based, 40 hours a week
Where
Remote
Field
Life, Physical, and Social Science
Project name
Neon
Posted
9/9/2026
Places left
30

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.