$30–50/hr
The role in one line
You design rigorous evaluation tasks for AI systems using real life sciences research materials to test complex experimental and analytical reasoning.
Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on micro1.
What you would do
- Develop complex, authentic research evaluation tasks including experimental design critique, protocol analysis, and data interpretation scenarios
- Source and synthesize research materials including raw data, lab protocols, literature, and technical documents
- Establish the appropriate scientific approach and biological conclusions for every evaluation scenario
- Develop detailed assessment rubrics focusing on methodological rigor, analytical reasoning, and information quality
- Build tasks that capture the complexity, uncertainty, and sophistication inherent in genuine research
Who they are looking for
- Degree in biology, life sciences, biochemistry, or related field
- At least 2 years conducting bench research or lab work involving experimental design and data analysis
- Proven skill in executing rigorous, multi-component scientific procedures and assessing technical soundness
- Mastery of scientific reasoning, source synthesis, and identifying methodological rigor issues
- Excellent communication ability, particularly in authoring and critiquing technical materials
Skills this role asks for
What the interview is likely to probe
1.Authentic task design from real data
Creating realistic scenarios from actual research materials ensures AI systems learn to handle genuine scientific complexity and ambiguity.
Expect something like: “You have primary data from a flawed enzyme kinetics experiment with noisy measurements and outliers. Would you simplify the data to make correct answers obvious, or preserve ambiguity to test AI reasoning? How would you construct this?”
2.Methodological evaluation and critique
Identifying and articulating methodological issues requires deep understanding of what makes research sound or fundamentally flawed.
Expect something like: “A protocol uses an inappropriate control group and lacks statistical power analysis. How would you structure a rubric item testing whether an AI recognizes both issues and explains why each matters?”
3.Source integration across formats
Real research synthesis requires combining information from diverse sources including tables, figures, text, and supplementary materials.
Expect something like: “You have a paper with conflicting conclusions between text and supplementary data tables. How would you design a task testing AI interpretation of this discrepancy?”
4.Rubric fidelity and objectivity
Precise rubrics ensure AI evaluation is consistent and scientifically defensible rather than subjective or ambiguous.
Expect something like: “How would you distinguish in a rubric between 'partially correct interpretation' and 'fundamentally misunderstands the data' for a multi-step analysis?”
5.Balancing realism and feasibility
Tasks must be realistic enough to test genuine reasoning but scoped appropriately for practical AI evaluation.
Expect something like: “An experiment involves 12 measured variables and multiple treatment groups. Would you present all data or curate a subset? How would you justify your choice?”
Exercise you may get
Design a complete evaluation task: source primary data from a research paper with a methodological issue, create a scenario requiring issue identification, author 3-4 rubric items with example responses, and write the correct answer with scientific rationale.
How to prepare
- Review recent life sciences literature to identify authentic research scenarios with methodological nuance and real-world complexity
- Study common experimental design flaws and data interpretation errors in your research domain
- Practice developing rubric items that distinguish surface-level answers from deep scientific reasoning
- Familiarize yourself with structuring multi-step scientific reasoning tasks for rigorous AI evaluation
Facts
- Pay
- $30–50/hr
- Eligible locations
- Bangladesh, Hong Kong, India, Indonesia, Japan, Kazakhstan, Kyrgyzstan, Malaysia, Pakistan, Philippines, Singapore, Sri Lanka, Taiwan, Thailand, Uzbekistan, Vietnam, Austria, Belarus, Belgium, Denmark, France, Germany, Greece, Italy, Netherlands, Portugal, Russia, Spain, Switzerland, United Kingdom, Argentina, Brazil, Chile, Colombia, Mexico, Peru, Algeria, Bahrain, Egypt, Iraq, Jordan, Kuwait, Lebanon, Libya, Morocco, Oman, Palestine, Qatar, Saudi Arabia, Tunisia, United Arab Emirates, United States, Canada, Nigeria, Kenya, South Africa, Ghana, Ethiopia
- Domain
- Sciences Research
- Role type
- generalist
- Posted
- 9/16/2026
- Open slots
- 50