Training Turk

Mercor listing

Task Writer: Project Agentic-MME

$60/hr

The role in one line

Design image-based test tasks that evaluate frontier AI models on investigative reasoning and information verification.

Written by Training Turk from the public listing; it may be incomplete or out of date. Read the full posting on Mercor.

What you would do

  • Identify high-resolution images where key details are small, degraded, or require manipulation to read
  • Write realistic prompts that necessitate image operations plus at least one web search
  • Record the shortest solution path with exact image operations used
  • Cite credible sources for every factual claim in the golden answer
  • Define answer format and document accepted response variants

Who they are looking for

  • Advanced degree holder (Master's, PhD, or equivalent professional credential)
  • Strong attention to detail across high-resolution image work and documentation
  • Excellent writing skills for clear prompt creation and task specifications
  • Skill at web research with habit of using primary, authoritative sources
  • Experience working on rubrics or task specifications preferred

Skills this role asks for

image analysis and manipulationai model evaluationprompt engineeringweb research expertiseprimary source identificationtask design and specificationattention to detailgolden answer documentation

What the interview is likely to probe

  1. 1.Image investigation task design

    The interviewer verifies you can create test images where models must actively manipulate visuals rather than passively interpret obvious details.

    Expect something like: “You've found an image with a small detail that's crucial to answering a question. Describe the image, explain why simple viewing isn't sufficient, and outline what manipulations a model would need to perform.”

  2. 2.Multimodal reasoning chain design

    Effective test tasks require integration of image operations with web research; weak prompts allow single-mode reasoning to succeed.

    Expect something like: “Design a prompt where a model must crop an image to find a product identifier, then web-search to retrieve information, then compare the results to original metadata. Describe the task and why naive approaches would fail.”

  3. 3.Primary source research quality

    Golden answers must cite authoritative sources; using secondary sources or common knowledge undermines task validity.

    Expect something like: “You need to verify a historical fact related to your test image. Describe your research process, the sources you'd check first, and how you'd validate information against multiple authoritative references.”

  4. 4.Solution documentation completeness

    Clear documentation of the shortest solution path ensures consistency in grading and helps identify where models diverge from optimal reasoning.

    Expect something like: “Document the solution to a task you've designed: what image operations does a model need to perform, in what sequence, and where would you expect the model to search for information?”

Exercise you may get

Design one complete test task: create or describe a high-resolution image with a small detail, write a realistic prompt requiring image manipulation and web search, and document the solution path with sources.

How to prepare

  • Study examples of how frontier models handle image manipulation and multi-step reasoning
  • Practice identifying and utilizing authoritative primary sources across multiple domains
  • Review how rubric-based task specification works in AI evaluation
  • Prepare examples of images where critical details are small or require enhancement

Facts

Pay
$60/hr
Commitment
hourly
Hours
20 per week
Work arrangement
remote · Remote
Domain
Miscellaneous
Posted
10/2/2026
Open slots
3