$65–75/task · Mercor · Task based
You red-team AI models on explosives and materials science to ensure they correctly handle the boundary between legitimate professional questions and dangerous misuse.
What you would do
- Write single-turn prompts testing model judgment on materials misuse potential at benign, dual-use, and adversarial levels
- Evaluate AI responses against defined safety policies and technical accuracy standards
- Author detailed reference answers explaining why correct responses handle the dual-use boundary appropriately
- Assess whether models refuse genuinely dangerous requests while fully answering legitimate professional questions
Who they want
- Federally certified explosives specialist or equivalent with forensic scene analysis and materials identification experience
- Background in forensic laboratory residue identification or post-incident analysis and reporting
- Experience with regulatory inspection of licensees, storage compliance, and diversion detection
- Strong technical writing skills, published research, or expert witness testimony experience
- Deep understanding of why the same knowledge serves both detection and evasion in dual-use contexts
Main skills
What the interview asks about
1.Forensic versus synthesis judgment
You must confidently distinguish when a materials question serves legitimate forensic case-building versus when it seeks information applicable to creation or concealment.
For example: “A question asks how to identify post-blast residue composition and match it against comparison standards from a suspected source. Is this legitimate casework or a probe for harmful knowledge?”
2.Regulatory compliance framing
Understanding what licensed professionals are actually permitted to do in their day-to-day work informs whether a model's refusal blocks routine legitimate operations.
For example: “A researcher at an approved testing facility asks about handling protocols for perchlorate-based mixtures during quality assessment. Classify this prompt's risk level and justify your judgment.”
3.Detection versus avoidance paradox
The same forensic signature that identifies a substance can also describe synthesis; regulators spot diversion through patterns that also reveal evasion methods.
For example: “A prompt describes specific diversion indicators used in regulatory inspections-the exact signals that reveal when materials are going missing. How should a model treat this question?”
4.Reference answer technical depth
Your explanations must convince non-specialists why a model refusal was justified or why an answer was appropriate, grounding the decision in actual forensic practice.
For example: “Draft a reference answer for a model that correctly refused a synthesis-focused prompt on a dual-use precursor. What technical reasoning justifies the refusal without being circular?”
A task you may get
Review an AI model response to a materials or explosives question, classify it as correct or incorrect, and draft a reference answer that explains the safety reasoning for non-specialist reviewers.
How to prepare
- Identify three real forensic or regulatory scenarios involving materials and map how each would be classified as benign, dual-use, or adversarial
- Write a sample reference answer on a materials topic, explicitly explaining the reasoning for holding or crossing the professional-dangerous boundary
- Review current red-teaming approaches in sensitive technical domains to understand how evaluation rubrics calibrate risk levels
- Sketch the workflow a forensic analyst and a regulatory inspector actually follow when handling materials, to ground prompt authenticity
The facts
- Pay
- $65–75/task
- Hours
- Task based
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Project name
- Neon
- Posted
- 9/14/2026
- Places left
- 100
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.