$65–75/task · Mercor · Task based
A nuclear security specialist who writes test prompts and judges AI responses on sensitive topics.
What you would do
- Write single-turn prompts at three threat levels to evaluate model judgment
- Score responses against a policy standard, documenting whether each was handled correctly
- Compose reference answers explaining the technical reasoning behind correct responses
- Identify scenarios where the line between legitimate and dangerous blurs
Who they want
- Deep background in material control, physical protection, or vulnerability assessment at licensed facilities
- Published research, technical writing, or expert witness experience
- Comfort discussing misuse scenarios without drawing on classified or NDA-restricted information
- Ability to work independently on writing-intensive, technical tasks
Main skills
What the interview asks about
1.Boundary judgment in safeguards
Real nuclear work requires distinguishing routine inventory questions from inquiries fishing for weaknesses, and this role demands you write both.
For example: “A model is asked about calculation methods for detecting material loss in a facility's accounting system. When is this a legitimate inspector question versus one seeking to hide a diversion?”
2.Dual-use response evaluation
Models must answer technical questions fully while refusing genuinely risky ones; you assess whether they held this line correctly.
For example: “A model provided a detailed explanation of vulnerability assessment methodology but declined to identify specific weak points at a real facility. Did it respond correctly?”
3.Technical explanation for lay reviewers
Your written rationales must let non-nuclear staff understand why a response was safe or unsafe.
For example: “Explain to a program manager with no nuclear background why a model's refusal of a question about sensor performance at a particular site was appropriate.”
4.Red-teaming craftsmanship
Prompts sitting on the policy line must sound natural and plausible, not obviously malicious, to test model robustness.
For example: “Write a prompt about insider threat assessment that a security officer at a facility would legitimately ask, then identify how an adversary might misuse the response.”
5.Scenario identification from context
You flag when a seemingly routine question carries hidden risk based on what a facility actually does.
For example: “A contractor asks about delay times for barrier systems at an unspecified location. Why does the context of their employment history matter to your threat classification?”
A task you may get
Evaluate three sample AI responses to nuclear security prompts: one incorrectly refused a legitimate question, one over-answered a risky one, and one held the right line. Explain each verdict.
How to prepare
- Review published nuclear security policy and inspector guidance to calibrate risk judgment
- Study examples of dual-use research in your domain to recognize escalation patterns
- Practice writing technical justifications that a non-specialist can follow
- Gather any prior red-teaming samples or technical writing to reference in discussions
The facts
- Pay
- $65–75/task
- Hours
- Task based
- Where
- Remote
- Field
- Other Engineering
- Project name
- Neon
- Posted
- 9/14/2026
- Places left
- 100
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.