$40–90/hr · micro1
PhD-level researcher or specialist who creates rigorous evaluation questions to assess AI model understanding of advanced research domains.
What you would do
- Design original, challenging questions requiring advanced reasoning and synthesis across multiple sources
- Source authoritative answers using primary literature with complete citations and documentation
- Test questions against AI systems to verify they have sufficient discriminatory power
- Iteratively refine questions to increase complexity while maintaining accuracy and preventing shortcuts
- Write all deliverables with high precision and clarity, eliminating ambiguity from technical language
Who they want
- PhD or active PhD candidacy in your field, or equivalent research experience as researcher or professor
- Strong track record of scholarly research and familiarity with primary literature in your domain
- Exceptional analytical thinking and attention to detail with strong written communication in English
- Demonstrated ability to source and triangulate information from authoritative references
- Self-directed professional who delivers expert-level work reliably in remote independent engagement
Main skills
What the interview asks about
1.Calibrating question difficulty for AI evaluation
Evaluation material must distinguish capable from incapable reasoning; interviewers check whether you understand what makes questions suitably discriminating.
For example: “You created a molecular biology question expecting moderate difficulty. An AI answered it correctly with confident reasoning in under a minute. What would you change? What difficulty level is appropriate for your field?”
2.Identifying and using authoritative sources
AI evaluation requires grounding in genuine expert knowledge, not fabricated or outdated material; interviewers verify you know how to distinguish authoritative from weak sources.
For example: “You're writing Q&A pairs on recent developments in machine learning interpretability. How would you determine what counts as an authoritative source, and what would you do if two major papers disagree about a finding?”
3.Preventing shortcut solutions in question design
Questions that allow pattern-matching without reasoning fail to evaluate genuine understanding; interviewers assess whether you can anticipate and block these gaps.
For example: “You drafted a question about statistical methodology. On review, someone notes it could be answered by recognizing keywords without actually understanding the concepts. How would you restructure it to force genuine analytical thinking?”
4.Technical precision in written communication
Ambiguous questions produce unreliable AI evaluations; interviewers check whether you balance technical accuracy with clarity suitable for evaluation.
For example: “Your answer to a neuroscience question is technically precise but uses five specialized terms that are standard in your subfield but may confuse outsiders. How would you decide whether to modify the language, add context, or keep it as written?”
5.Synthesizing across disciplinary boundaries
Frontier research increasingly requires integrating concepts from multiple fields; interviewers verify you can design questions that test genuine synthesis.
For example: “You're creating a question that requires understanding both statistical learning theory and biological plausibility constraints. How would you structure this to test synthesis rather than just isolated knowledge of each component?”
A task you may get
Create one original question-answer pair at the frontier of your field that tests advanced reasoning or methodological understanding, then explain why this question effectively evaluates AI capability and how you would verify the answer's accuracy.
How to prepare
- Review the most recent papers published in your field to identify emerging questions that push current boundaries
- Document the authoritative sources in your specific area and understand their citation relationships and credibility hierarchy
- Study how AI systems currently approach questions in your domain and anticipate which shortcut solutions they might employ
- Refine your ability to write technical explanations that maintain precision while remaining clear to researchers from adjacent disciplines
The facts
- Pay
- $40–90/hr
- Open to
- Bangladesh, Hong Kong, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestine, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia
- Field
- Sciences Research
- Role type
- Generalist
- Posted
- 8/21/2026
- Places left
- 5
We wrote this page from the public micro1 listing. It may be out of date, so read the full posting before you apply.