Training Turk

Drug Discovery Scientist

$60–90/hr · Mercor · Hourly, 20 hours a week

You design challenging drug discovery problems that train AI to reason authentically about pharmaceutical research by engineering realistic scenarios with hidden answers in messy evidence.

What you would do

  • Create realistic drug discovery research challenges where the correct answer exists in provided evidence but cannot be found through simple lookup
  • Assemble data rooms from real research sources ensuring every key fact traces back to an original publication or database
  • Design grading rubrics that assess whether AI reasoning is scientifically sound and well-supported, regardless of method or presentation
  • Validate and calibrate problems to ensure they require appropriate levels of reasoning without being unsolvable or trivially easy
  • Revise problem design and grading criteria based on reviewer feedback to meet quality and difficulty standards

Who they want

  • PhD or equivalent track record in enzyme mechanistic chemistry, peptide/antibody engineering, fragment discovery, or medicinal chemistry
  • Proven ability to pose problems whose solutions can be defended rigorously even when other solutions might seem plausible
  • Expertise evaluating reasoning quality and distinguishing strong arguments from weak ones, not merely checking correctness of final answers
  • Minimum 20 hours per week commitment, with flexibility to scale to 40 hours per week as needed
  • Access to Claude Code (Max subscription required; reimbursement provided) and comfort working in browser-based studio environments

Main skills

Mechanistic enzymologyStructure guided peptide designAntibody engineering

What the interview asks about

  1. 1.Problem design and difficulty calibration

    AI training quality depends on problems that force genuine reasoning; evaluators must confirm candidates can design appropriately difficult, solvable challenges.

    For example: “Design a fragment-based discovery problem where the AI chooses between chemical series based on binding data and selectivity. How would you structure evidence so the answer requires reasoning?”

  2. 2.Real data sourcing and messiness

    AI trains best on authentic, imperfect data; evaluators assess whether candidates will assemble realistic data rooms or create idealized scenarios.

    For example: “You have identified three conflicting binding affinity measurements for the same compound from different labs and different assay conditions. How would you include these in your data room, and what guidance would you provide about resolving conflicts?”

  3. 3.Grading rubric design for reasoning

    Rubrics must evaluate thinking, not just answers; evaluators check whether candidates can articulate what makes reasoning sound versus weak.

    For example: “An AI proposes skipping a particular experiment and moving directly to in vivo studies. Walk through how you would score this recommendation - what reasoning would make it acceptable, and what reasoning would make it unacceptable?”

  4. 4.Problem validation and adjustment

    Candidates must test whether their problem actually works; evaluators want to know if they will blindly submit or validate and adjust based on outcomes.

    For example: “Your problem validation shows a baseline approach solves it easily, suggesting difficulty is too low. How would you decide: reformulate the problem or accept the easier difficulty?”

  5. 5.Domain expertise and molecular judgment

    Only candidates with deep domain knowledge can design authentic problems; evaluators assess whether candidates understand their specialty's constraints and trade-offs.

    For example: “You are designing a peptide engineering problem where candidates must balance potency against a known off-target liability. What evidence would you need to include so the reasoning is authentic to how medicinal chemists actually approach this trade-off?”

A task you may get

Design a complete drug discovery problem including: the research prompt with a specific challenge, a data room with 3-5 real sources, and a grading rubric with 5-7 criteria; explain why each problem element requires genuine reasoning.

How to prepare

  • Identify a recent paper in your field and reverse-engineer what evidence would be needed for an AI to reach the paper's conclusions independently
  • Review a complex design decision from your career and practice explaining why your approach was best among plausible alternatives
  • Gather 3-4 real papers from your specialty and practice constructing realistic scenarios around them where critical evidence comes from different sources

The facts

Pay
$60–90/hr
Hours
Hourly, 20 hours a week
Where
Remote
Field
Life, Physical, and Social Science
Posted
8/31/2026
Places left
20

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.