Training Turk

Emergency Response Specialist

$65–75/task · Mercor · Task based

A radiological safety expert who writes and evaluates test prompts for frontier AI models to ensure they correctly judge misuse potential in your technical domain.

What you would do

  • Craft single-turn prompts across three levels that test whether an AI model can tell a routine professional question from one seeking dual-use or adversarial information
  • Evaluate the AI's responses against a defined policy standard, documenting whether each response was handled correctly
  • Write reference answers and technical reasoning showing what a correct response looks like and why
  • Identify edge cases where the same concept (isotopes, shielding, dispersal models) legitimately serves both professional and harmful purposes

Who they want

  • Field experience with radiological event response, emergency planning at licensed facilities, or radiological consequence modeling
  • Demonstrable ability to write technical explanations that non-specialists can follow, ideally shown through published research or expert materials
  • Judgment to recognize when a request is fishing for dangerous information despite using professional language
  • Comfort writing about misuse scenarios for sustained periods while understanding the importance of this red-teaming work

Main skills

Radiological safetyRed teaming AI modelsDual use detection

What the interview asks about

  1. 1.Drawing the safety line

    You must articulate how the same technical concept can serve legitimate and harmful purposes, and explain precisely where an AI should accept or refuse. This judgment comes only from real experience in the field.

    For example: “A model gets asked about shielding calculations for industrial radiography. How do you evaluate whether its response is appropriately detailed without crossing into dangerous territory? Write a reference answer.”

  2. 2.Recognizing dual-use requests

    You write prompts that sit exactly on the line between professional inquiry and adversarial probing. This requires understanding how practitioners actually phrase questions versus how someone seeking dangerous information would phrase them.

    For example: “You're drafting a dispersal modeling prompt for someone claiming to manage facility evacuation. Which details make it realistic for the AI but which do you omit? How do you know when you've gone too far?”

  3. 3.Technical documentation under scrutiny

    Non-specialists will read your evaluations to judge whether the model held the line. Your written rationale must be clear and defensible, showing the technical reasoning without assuming specialized knowledge.

    For example: “You're explaining to a non-radiologist why a model correctly refused to calculate dose rates for an unshielded source in an urban area, or why it should have. What do you include in your rationale to make the reasoning clear and persuasive?”

  4. 4.Incident-specific prompt realism

    Your prompts need the texture of real radiological work to properly test an AI. This comes from having seen how responders, planners, or field teams actually phrase requests in emergency or operational contexts.

    For example: “You've run a full-scale radiological exercise. Now write a benign prompt about medical countermeasures grounded in realistic operational details. What makes it authentic to the AI?”

  5. 5.Identifying model weaknesses

    Your evaluation should spot cases where the model is inconsistent, overly cautious on routine questions, or not cautious enough. These gaps reveal where the model's safety line is fuzzy and where red-teaming coverage matters most.

    For example: “The model accepts dose calculations from a medical physicist but refuses a nearly identical question about dispersal monitoring. Do you flag this as a failure, reasonable distinction, or training gap?”

  6. 6.Sustaining focus on sensitive material

    You work with misuse scenarios and sensitive technical details over extended periods. This demands both intellectual stamina and a clear understanding of why this red-teaming work matters and where your ethical boundaries are.

    For example: “After several hours of writing adversarial scenarios in your domain, you notice fatigue affecting your judgment of what's dangerous. How do you recalibrate? What's your process for stepping back and ensuring you're still drawing the line reliably?”

A task you may get

Write a benign prompt and a dual-use prompt about a specific radiological concept you've encountered in your work, then draft a reference answer showing how an AI should respond to each. Explain your reasoning for drawing that distinction.

How to prepare

  • Bring a sample of your technical writing, published research, or expert documentation that demonstrates how you explain radiological safety to varied audiences
  • Prepare an example of a real decision you made in radiological emergency planning or field response where distinguishing routine from dangerous was critical
  • Review recent AI safety literature on red-teaming and dual-use research to understand the broader context and language around this work
  • Reflect on your experience with radiological exercises or events: which scenarios most sharpened your judgment about where safety lines sit?

The facts

Pay
$65–75/task
Hours
Task based
Where
Remote
Field
Life, Physical, and Social Science
Project name
Neon
Posted
9/14/2026
Places left
100

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.