$23–27/hr · Mercor · Part time, 7 hours a week
Bilingual PhD scientist reviewing AI model responses on sensitive technical topics to ensure scientific accuracy and responsible handling of potentially dangerous information.
What you would do
- Write sophisticated scientific questions in Hindi that test model reasoning across chemistry, biology, radiological, or nuclear subject matter
- Evaluate whether AI responses are scientifically accurate, complete, and appropriately cautious about hazards
- Assess how well models handle sensitive dual-use information without unnecessary scaremongering or dangerous oversimplification
- Apply classification frameworks to categorize conversations and identify safety concerns in real time
- Provide structured feedback that helps improve how AI systems communicate technical knowledge responsibly
Who they want
- PhD in chemistry, biology, or closely related science field, ongoing or completed
- Native or near-native Hindi fluency combined with business-level written English
- Deep practical knowledge of laboratory and computational techniques in your specialty
- Sound judgment about scientific safety, dual-use risks, and responsible information handling
- Experience in specialized subfields like organic synthesis, molecular biology, radiochemistry, or virology preferred
Main skills
What the interview asks about
1.Crafting technical questions in Hindi
You need to demonstrate scientific depth in Hindi, not just translation skills. Your prompts must be sophisticated enough to genuinely test what the model knows.
For example: “Write a Hindi question asking the AI to explain nucleophilic aromatic substitution. What chemistry terms ensure precision without losing clarity for native Hindi speakers?”
2.Evaluating scientific accuracy without AI expertise
You're hired for your science knowledge, not ML background. The interviewer wants to know you can spot when a response sounds plausible but contains subtle technical errors.
For example: “An AI claims crystalline copper sulfate suits pH measurement in biological samples. Explain why this is wrong and how you'd flag the solubility or reactivity gap in your review.”
3.Recognizing dual-use knowledge risks
The same information that helps a legitimate researcher can enable harm. You need to distinguish between responsible science communication and dangerous guidance.
For example: “A Hindi question asks how to isolate a particular bacterium from soil samples. The model provides accurate, detailed culturing instructions. When would this be helpful scientific guidance versus concerning, and how does context matter in your safety judgment?”
4.Domain-specific safety judgment
Chemistry, biology, and nuclear domains have different risk profiles. Your specialized knowledge of what's genuinely dangerous in your field is irreplaceable.
For example: “You're reviewing model responses about radiochemistry. The system accurately describes the behavior of Technetium-99m but doesn't mention gamma-ray hazards or shielding requirements. Is this an oversight, a safety gap, or appropriate for the question level?”
5.Translating safety judgment across languages
Danger isn't language-specific, but how you communicate it is. You need to evaluate safety in Hindi while ensuring your feedback works for English-speaking teams.
For example: “A Hindi question describes symptoms of chemical exposure. The model responds in Hindi using technical terms that translate misleadingly in English. Walk me through how you'd flag this - is it a translation problem, scientific inaccuracy, or safety concern?”
A task you may get
Review a Hindi prompt-response about organic synthesis. Assess accuracy, missing safety info, technical depth, and dual-use risks for your specialty.
How to prepare
- Identify 2-3 technical papers or recent textbooks in your specialty and review how safety is discussed in scientific literature
- Prepare examples of Hindi scientific terminology in your field - ensure you can communicate complex ideas clearly in both languages
- Research current discussions about responsible AI in scientific domains and dual-use information governance
- Think about the most dangerous misconception someone could have from bad technical information in your field, and how you'd recognize it in model responses
The facts
- Pay
- $23–27/hr
- Hours
- Part time, 7 hours a week
- Where
- Remote · Remote — South Asia preferred
- Posted
- 9/4/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.