$45–60/hr · Mercor · Part time, 40 hours a week
Humanities scholar who grades AI-generated text and writes reference answers to help train models in historical, philosophical, and literary interpretation.
What you would do
- Write discipline-specific prompts at multiple levels of difficulty that test AI understanding of humanities concepts
- Evaluate AI-generated responses for accuracy, source interpretation, and scholarly reasoning quality
- Create reference answers explaining what correct scholarly analysis looks like and why
- Identify fluent-sounding but incorrect AI responses that misread texts, flatten debates, or fabricate citations
- Write clear rationales explaining the reasoning behind each evaluation decision
Who they want
- Graduate degree in humanities field (history, classics, philosophy, literature, religious studies, cultural studies, or equivalent)
- Ability to clearly articulate scholarly judgment in writing with strong logical reasoning
- Comfort handling ambiguous work and capacity to raise issues with vague guidance
- Experience with academic writing, published research, or expert testimony demonstrating scholarly communication skills
- Awareness of limitations and willingness to pause or step away from misuse-related content without penalty
Main skills
What the interview asks about
1.Textual Interpretation Verification
Humanities experts must catch cases where AI produces grammatically fluent responses that fundamentally misread or misrepresent what primary sources actually say, preventing trained models from embedding interpretive errors.
For example: “An AI response interprets a 19th-century philosopher's argument as supporting a position the philosopher actually opposed. How would you identify this error and explain the correct interpretation in your reference answer?”
2.Scholarly Debate Recognition
AI must understand that scholarly fields involve genuine debate between legitimate positions; experts distinguish correct minority views from incorrect claims and help train models to engage appropriately with scholarly disagreement.
For example: “In literary criticism, scholars genuinely disagree about whether a work's ending suggests redemption or despair. How would you design a prompt and reference answer that helps AI recognize this as legitimate scholarly debate rather than factual certainty?”
3.Citation Fabrication Detection
AI models can generate plausible-sounding false citations; experts must flag fabricated sources, misattributed quotes, and invented scholarly work to prevent models from learning to cite nonexistent authorities.
For example: “An AI response quotes a scholar's exact words from a work you know doesn't exist. How would you document this fabrication and use it to help train AI to verify sources before citing them?”
4.Historical Context Accuracy
Humanities interpretation depends on accurate historical context; experts assess whether AI places ideas, events, and cultural phenomena in correct temporal and social frameworks.
For example: “A response explains a medieval religious text using modern concepts of psychology and individual rights that didn't exist in that era. How would you revise the reference answer to ground interpretation in actual period context?”
5.Argument Structure and Logic
Scholarly reasoning requires sound logical structure; experts evaluate whether AI arguments support their claims, avoid non sequiturs, and build coherent interpretations rather than disconnected observations.
For example: “An AI response about a historical event presents three observations that don't logically support the stated conclusion. How would you explain what's missing and model proper historical reasoning?”
A task you may get
Write a prompt in your field testing AI understanding of either a primary text interpretation or a scholarly debate. Provide both a strong reference answer and a flawed response, then explain why the flawed response fails the standard you've set.
How to prepare
- Identify three recent scholarly papers in your field and think about what AI needs to understand to engage with them correctly
- Review examples of common historical, philosophical, or literary misinterpretations in popular writing to recognize patterns of error
- Draft 2-3 prompts in your field that sit at the edge of what AI should handle correctly versus refuse appropriately
- Prepare a brief sample of your academic writing demonstrating scholarly communication clarity
The facts
- Pay
- $45–60/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Humanities
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.