Training Turk

English Language and Literature Expert

$50/hr · Mercor · Hourly

You evaluate language model performance on English tasks and provide expert feedback to improve their linguistic accuracy and naturalness.

What you would do

  • Assess LLM-generated English across tasks like summarization, dialogue, and grammar correction
  • Annotate text for grammatical, syntactic, semantic, and stylistic properties
  • Edit and rewrite model outputs to meet linguistic and cultural standards
  • Identify specific model errors and linguistic gaps in performance
  • Write detailed analysis explaining how and why outputs fall short

Who they want

  • Native or near-native English fluency with command of grammar, syntax, vocabulary, and style
  • Professional experience in editing, proofreading, writing, or linguistic annotation work
  • Deep understanding of linguistic nuances, regional variations, and context-appropriate English usage
  • Skill in spotting subtle grammar mistakes, stilted phrasing, and unnatural language patterns
  • Strong analytical judgment and written communication skills

Main skills

LLM response evaluationLinguistic annotation and analysisGrammar and syntax error detection

What the interview asks about

  1. 1.Subtle error detection

    Language models commonly produce grammatically correct but linguistically odd output; spotting these errors before training determines whether the model learns to sound natural.

    For example: “A model writes: 'The findings suggest that increased funding would be beneficent to the program's outcomes.' Native speakers rarely say 'beneficent' in this register. What specific issue would you note, and why does this matter in training data?”

  2. 2.Context-appropriate revision

    The right rewrite depends on understanding audience, formality level, and intent; feedback that ignores context teaches the model to produce mismatched language.

    For example: “A model generates dialogue: 'I would advise you to reconsider this matter at your earliest convenience.' This sounds formal for a casual conversation between friends. How would you rewrite it, and what linguistic principle guides your choice?”

  3. 3.Ambiguity and coherence judgment

    Models struggle with reference resolution and coherence across sentences; identifying where meaning breaks down teaches clearer reasoning patterns.

    For example: “A summary says: 'The policy aims to reduce emissions. It should be stricter.' The second sentence has unclear reference: does 'it' mean the policy or the reduction? How would you flag and fix this?”

  4. 4.Dialect and register knowledge

    English varies hugely by region and social context; models need training data that reflects this variation to avoid producing oddly formal or inappropriate language.

    For example: “A model generates customer service text: 'We ain't got that in stock right now, but we can order it for you.' The contraction feels mismatched to professional context. What would you change and why?”

  5. 5.Reasoning error identification

    Models sometimes produce logical gaps or inconsistencies that sound natural but fail on content; annotators must spot where reasoning breaks down.

    For example: “On a task asking for logical deduction, a model writes: 'All birds can fly. Penguins are birds. Therefore, penguins can fly.' What error would you flag in your feedback, and how would you explain why this matters?”

  6. 6.Stylistic consistency across length

    Long outputs are prone to shifts in tone or register; your feedback helps models maintain appropriate style throughout generated text.

    For example: “A 500-word article starts in formal academic register but slips into casual language halfway through. How would you describe this inconsistency to an engineer for training purposes?”

A task you may get

Evaluate three short LLM outputs (summary, dialogue, reasoning task) for linguistic errors, annotate features, and provide a brief written analysis of each model failure with suggestions for improvement.

How to prepare

  • Review linguistic terminology you'll use to describe errors: distinguish between grammatical mistakes, stylistic awkwardness, and semantic problems
  • Read recent examples of LLM output in domains like summarization and dialogue to familiarize yourself with common failure patterns
  • Prepare a personal style guide noting register, formality, and voice preferences so your feedback is consistent and articulated clearly

The facts

Pay
$50/hr
Hours
Hourly
Where
Remote
Field
Language and Audio
Posted
6/2/2025
Places left
2

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.