$65–75/task · Mercor · Task based
An expert from chemical defense background who creates red-teaming prompts to test whether AI models correctly distinguish routine protective knowledge from dangerous chemical threat information.
What you would do
- Write challenging single-turn prompts at three risk levels: benign, dual-use, and adversarial in chemical defense domains
- Evaluate AI model responses and judge whether each was handled correctly per safety policy standards
- Write reference answers showing correct technical responses and explaining defensive versus hostile reasoning
- Calibrate prompts to sit exactly on the boundary between legitimate protective questions and dangerous information
- Assess whether AI correctly distinguishes countermeasure protocols from threat assessment procedures
Who they want
- Hands-on experience in chemical defense from institutional setting (government research, DoD, CDC, or allied agency)
- Deep understanding of countermeasures, protection, detection, and verification in chemical threat work
- Ability to tell routine professional questions from ones seeking hostile information
- Background in technical documentation, published studies, or expert witness testimony
- Prior AI evaluation or red-teaming experience preferred but not required
Main skills
What the interview asks about
1.Defensive knowledge calibration
A response that correctly answers a countermeasure question can look identical to one answering a threat assessment; context and intent determine safety.
For example: “Someone asks about toxicological effects and exposure dose-response for a specific agent. This is legitimate for protection but also useful for threat planning. How would you calibrate a dual-use prompt that tests whether the model can answer defensively?”
2.Protection versus harm information
Decontamination procedures protect responders but can also inform an attacker about agent properties; only practitioners know this boundary reliably.
For example: “Design a prompt about decontamination procedures that separates routine protective knowledge from information that would enable someone to plan a release.”
3.Medical countermeasure reasoning
Treatment protocols drive threat assessment; the same pharmacological knowledge serves both. Distinguishing contexts requires domain expertise.
For example: “Write a benign prompt about medical countermeasure development and an adversarial variant asking similar information but from a threat perspective. How would you construct these to sit exactly on the policy boundary?”
4.Policy evaluation and response assessment
Experts determine whether a model's response violated policy; this requires judgment about what information is defensible versus dangerous.
For example: “A model answers a question about CWC verification procedures in detail. Against a policy allowing defensive information but not procedures, evaluate whether this violates policy and explain your reasoning.”
5.Reference answer with defensive reasoning
Non-specialists need to understand why a technical answer was correct for defensive use; connecting knowledge to defensive purpose is key.
For example: “Write a reference answer for a prompt about chemical agent persistence. How would you explain the technical content and connect it specifically to defensive needs rather than threat potential?”
A task you may get
Create three prompts (benign, dual-use, adversarial) in your chemical defense specialty, evaluate provided model responses for each, and write one reference answer that explains technical correctness with explicit defensive reasoning.
How to prepare
- Reflect on a specific example from your work where legitimate protective knowledge could be misused without context
- Review existing red-teaming frameworks for chemical and biological threat evaluation
- Study how policy differentiates defensive versus offensive use of technical information
- Prepare an explanation of how you'd convey a complex chemical threat judgment to someone without defense background
The facts
- Pay
- $65–75/task
- Hours
- Task based
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Project name
- Neon
- Posted
- 9/14/2026
- Places left
- 100
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.