$65–75/task · Mercor · Task based
You red-team AI models on toxicology and forensics, writing challenging prompts and evaluating whether models correctly distinguish legitimate professional chemistry questions from requests enabling misuse.
What you would do
- Write single-turn prompts that test model judgment on dual-use chemistry topics at multiple difficulty levels
- Evaluate how models respond: whether they answer appropriate questions fully and refuse genuinely dangerous requests
- Assess responses against explicit misuse policy standards and provide pass/fail judgments with reasoning
- Write reference answers explaining the technical conclusion and the boundary judgment that drove the evaluation
- Document your reasoning in clear prose so reviewers and model trainers understand your judgment
Who they want
- Demonstrated casework experience identifying substances, precursors, and clandestine laboratory chemistry
- Deep knowledge of forensic toxicology, postmortem analysis, or chemical investigation
- Ability to write complex technical analysis clearly for non-specialists and expert audiences
- Experience with expert testimony, published research, or regulatory documentation demonstrating technical writing skill
- Willingness to spend sustained periods writing detailed analyses about misuse scenarios in your field
Main skills
What the interview asks about
1.Distinguishing routine casework from misuse fishing
The hardest evaluation cases look legitimate at first; you must recognize the subtle indicators showing whether a question genuinely seeks professional information or seeks to extract misuse instructions.
For example: “A question asks about precursor structure 'because I'm analyzing clandestine waste.' When is this legitimate casework, and what language signals synthetic intent?”
2.Writing prompts that expose model judgment gaps
Obvious questions don't train useful policy adherence; your prompts must force models to think through ambiguous cases and gray zones.
For example: “Compose a prompt involving forensic toxicology that seems benign on the surface but contains subtle dual-use indicators. Explain why this would be a good evaluation task.”
3.Documenting boundary reasoning clearly
Your evaluation is the ground truth for training; reviewers and model trainers must understand your judgment, not just see pass/fail.
For example: “You evaluated a model response as 'insufficient refusal' on a question about precursor identification. Write the brief explanation you'd provide explaining why the response crossed the policy boundary.”
4.Identifying when you need to escalate
Some requests go beyond your specialty or raise flags requiring coordinator input; recognizing when a prompt itself is inappropriate is part of red-teaming.
For example: “A proposed prompt asks you to evaluate a model's response to a question clearly involving export-controlled chemical synthesis data. How would you flag this for the red-team coordinator?”
How to prepare
- Prepare 2-3 real casework examples from your practice distinguishing legitimate from illegitimate inquiry
- Review your specialty's standards on expert testimony and documentation of chemical findings
- Collect examples of dual-use chemical knowledge you've navigated professionally - times you answered a question carefully or declined to answer
- Prepare brief written case summaries showing your reasoning when a question sat on the benign/misuse boundary
The facts
- Pay
- $65–75/task
- Hours
- Task based
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Project name
- Neon
- Posted
- 9/14/2026
- Places left
- 100
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.