$65–75/task · Mercor · Task based, 40 hours a week
An expert who creates test prompts for AI models to evaluate whether they correctly judge chemistry misuse risk while answering legitimate questions.
What you would do
- Write challenging single-turn prompts at three risk levels: benign, dual-use, and adversarial
- Evaluate AI model responses and judge whether each was handled correctly per policy standards
- Write reference answers showing correct technical responses and explaining reasoning
- Calibrate prompts to sit on the exact boundary between routine questions and dangerous requests
- Assess AI performance across synthesis, analytical chemistry, and safety domains
Who they want
- Deep hands-on working experience in chemical synthesis, analytical methods, or process safety contexts
- Understanding of where theoretical knowledge ends and procedural knowledge begins
- Prior work in technical documentation, peer-reviewed publications, or expert testimony
- Prior AI evaluation or red-teaming experience preferred but not required
- Ability to explain complex domain judgments clearly to non-specialists
Main skills
What the interview asks about
1.Calibration of dual-use prompts
A model that wrongly refuses a routine question is as valuable a failure case as one that answers something it shouldn't; distinguishing these is core expertise.
For example: “You're creating a dual-use prompt about protective equipment for a specific threat. What makes this question routine versus fishing for information an attacker would use? How would you label and justify your calibration?”
2.Mechanism versus procedure distinction
The danger line sits between understanding how a reaction works and knowing the specific steps to perform it; only practitioners reliably know this boundary.
For example: “Someone asks about precursor chemistry and synthesis intermediates. Which aspects are legitimate analytical knowledge, and which enter adversarial territory? How would you write prompts that test this boundary?”
3.Evaluation against safety policy
Policies define what models should refuse or answer fully; expert judgment determines whether a specific response actually follows the policy.
For example: “A model answers a detection method question fully. Against a policy allowing defensive information but not procedures, is this response acceptable or policy-violating? What's your reasoning?”
4.Reference answer documentation
Non-specialist audiences need to understand why a technical decision was correct or incorrect without domain expertise.
For example: “Write a reference answer for a question about countermeasure development. How would you explain the technical correctness to someone without a chemistry background?”
5.Adversarial prompt construction
Effective adversarial prompts reveal where models actually draw risk boundaries; poorly designed prompts miss real vulnerabilities.
For example: “Design an adversarial prompt around toxicological assessment that tests whether a model confuses offensive assessment with defensive preparation. What makes it effectively adversarial?”
A task you may get
Create three prompts (benign, dual-use, adversarial) around a chemistry topic you know deeply, evaluate a provided AI response for each against a simple policy, and write one reference answer with technical reasoning.
How to prepare
- Review published red-teaming frameworks and how misuse risk is typically assessed
- Reflect on a complex technical boundary in your field where mechanism knowledge differs from procedural knowledge
- Study examples of policy violation versus compliance in your domain
- Prepare a brief example of how you'd explain a complex domain judgment to a non-specialist
The facts
- Pay
- $65–75/task
- Hours
- Task based, 40 hours a week
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Project name
- Neon
- Posted
- 9/9/2026
- Places left
- 30
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.