Training Turk

Energetic Materials Expert for Redteaming

$65–75/task · Mercor · Task based, 40 hours a week

You test whether AI models correctly judge the misuse risk of technical requests about explosives and energetic materials.

What you would do

  • Create prompts at benign, dual-use, and adversarial levels to test whether AI judges misuse risk correctly
  • Evaluate AI responses against policy standards, determining if refusals and acceptances are appropriate and safe
  • Write reference answers showing what correct responses look like and explaining technical reasoning
  • Score performance and provide rationales for evaluations in language non-specialists can understand
  • Document novel AI behaviors discovered during evaluation testing

Who they want

  • Deep expertise in one of: hazardous device assessment, pyrotechnics chemistry, propulsion systems engineering, incident forensics analysis, or commercial blasting
  • Track record handling sensitive materials professionally: demonstrated safety protocols, regulatory compliance, incident reporting or render-safe procedures
  • Prior experience in AI evaluation, red-teaming, or technical writing (published research, expert witness testimony, incident reports preferred)
  • Ability to write clearly for non-specialist audiences without losing technical precision
  • Comfort with sustained focus on misuse scenarios; can pause or step back without penalty

Main skills

Explosives chemistry and initiationPyrotechnics formulation and testingPropulsion systems design

What the interview asks about

  1. 1.Judging legitimate vs dangerous requests

    Field experts immediately see the difference a generalist misses; your judgment teaches AI what to allow and what to refuse.

    For example: “Two prompts: one asks how to calculate detonation velocity for a new formulation (a mining engineer's question), the other asks the same but frames it for 'hobby pyrotechnics'. Are they both legitimate, both dangerous, or different? Why?”

  2. 2.Writing adversarial prompts that test policy

    Generic 'how to build an explosive' is easy to refuse; testing the model's judgment requires prompts that are technically coherent and policy-adjacent.

    For example: “Design a dual-use prompt about propulsion chemistry that would be legitimate for a researcher but could enable misuse if the AI gives too much detail. How do you frame it?”

  3. 3.Explaining technical judgment to non-specialists

    Your evaluation only matters if risk officers, lawyers, and ML engineers can follow your reasoning without relearning your entire field.

    For example: “You refused an AI's response to a formulation question because it omitted a critical stability constraint. Write a 2-3 sentence explanation for a non-chemist about why this matters.”

  4. 4.Recognizing when AI fails asymmetrically

    If AI over-refuses legitimate questions or under-refuses dangerous ones, your documentation shapes how the model is retrained.

    For example: “The AI refused a routine question from a mining engineer about blast design but answered a carefully framed adversarial variant. What does this pattern tell you about the model's policy?”

  5. 5.Managing sustained focus on sensitive content

    Fatigue and emotional weight can degrade judgment; you must recognize when you need to pause and maintain rigor.

    For example: “After 8 hours of red-teaming on blast scenarios, you're evaluating a prompt that feels familiar but you can't remember which version you just rated. What's your move?”

A task you may get

Write one benign, one dual-use, and one adversarial prompt on a technical question in your field. For each, score how an AI should respond and write a brief rationale.

How to prepare

  • Collect 3-5 real technical inquiries from your field (from publications, forums, professional groups) and note which are routine vs risky
  • Write a 1-page example of technical risk reasoning: a legitimate question, why it's legitimate, and how someone might misuse similar information
  • Review sample red-teaming briefs or misuse policies to understand how policy translates to judgment calls

The facts

Pay
$65–75/task
Hours
Task based, 40 hours a week
Where
Remote
Field
Life, Physical, and Social Science
Project name
Neon
Posted
9/9/2026
Places left
41

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.