$65–75/task · Mercor · Task based
A chemistry expert writes test prompts for AI models to evaluate their judgment on energetic materials and propulsion safety.
What you would do
- Draft multi-level prompts that test an AI model's ability to handle dual-use information responsibly
- Evaluate model responses for technical accuracy and compliance with safety policies
- Write detailed reference solutions showing correct reasoning and explaining judgment calls
Who they want
- Formulation or synthesis background with energetic compounds in licensed or government facilities
- Direct experience with calorimetry, thermal stability testing, and sensitivity characterization methods
- Published research, technical writing samples, or expert witness testimony demonstrating communication ability
- Willingness to discuss misuse scenarios in your field for extended work periods
Main skills
What the interview asks about
1.Boundary recognition in dual-use
The interviewer needs confidence you can author prompts and evaluate answers where the same data serves legitimate and dangerous purposes - something generalists cannot reliably judge.
For example: “Model receives question on thermal decomposition profiles for specialty oxidizer. Data is used in production but hostile actor could repurpose it. Label this. Correct response?”
2.Technical reasoning assessment
Spotting code or logic that runs but is wrong, or answers that sound authoritative but miss key constraints, requires hands-on engineering judgment that cannot be faked.
For example: “An AI proposes a binder formulation that meets all stated burn-rate targets but ignores a known mechanical brittleness issue above a certain cure temperature. Your written explanation needs to justify why this fails.”
3.Explaining tacit knowledge clearly
Red-teaming requires translating field intuition into explicit criteria so others can apply your standards and improvements compound over time.
For example: “Explain to ML researcher why scale-up decisions depend on production experience factors that literature might miss. How translate field intuition to explicit criteria?”
4.Safety policy compliance evaluation
AI safety requires consistent application of policies across diverse scenarios; your judgment on when policies apply and how strictly is essential.
For example: “Model response follows your stated policy but cuts corners on one safety specification. How would you score this and what would you write to help the research team tighten the policy?”
5.Identifying subtly flawed technical reasoning
Catching plausible-sounding but incorrect technical reasoning requires hands-on domain expertise; this is where field experience is irreplaceable.
For example: “Model gives correct formulation chemistry but misses a storage stability issue under specified conditions. How would you document this gap and its practical consequence?”
A task you may get
Draft three prompts at increasing difficulty levels on a pyrotechnic or propellant topic, then score sample model outputs against a safety rubric you define.
How to prepare
- Review recent red-teaming frameworks used in AI safety to understand how your domain fits into broader misuse prevention
- Identify 2-3 recent papers or technical documents from your field where the same result has legitimate and sensitive applications
- Prepare one written example (memo, blog post, or deposition excerpt) showing your ability to explain materials science nuance to non-specialists
The facts
- Pay
- $65–75/task
- Hours
- Task based
- Where
- Remote
- Field
- Life, Physical, and Social Science
- Project name
- Neon
- Posted
- 9/14/2026
- Places left
- 100
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.