$45–65/hr · Mercor · Part time, 35 hours a week
Linguist or instructional designer translating program requirements into precise, consistent rater guidelines for AI training data across specialized domains.
What you would do
- Convert ambiguous specifications from finance, retail, insurance, legal, and sports into clear, actionable rating instructions
- Design and refine rubrics that raters can apply uniformly, including for edge cases and unusual scenarios
- Review and revise draft guidelines to eliminate contradictions, ambiguity, and coverage gaps
- Coordinate with domain specialists and research teams to verify accuracy and consistency
- Document your reasoning and obtain stakeholder approval before deployment
Who they want
- 3+ years writing or refining guidelines for human raters in GenAI or RLHF contexts
- Demonstrated ability to work across multiple specialized domains and translate nuance into clear language
- Strong track record of resolving specification ambiguity with concrete before-and-after examples
- Ability to commit 35 hours per week, primarily weekdays
- Excellent written communication and precision in explaining complex guidance
Main skills
What the interview asks about
1.Resolving contradictory requirements
Guidelines that contain internal conflicts lead raters to escalate constantly and produce inconsistent training labels; your ability to find and fix this determines output quality.
For example: “You received draft finance rater guidelines where the threshold for flagging high-risk trades conflicted with the mark-to-market rule. Walk me through how you identified the conflict and what you proposed to the program team.”
2.Edge case coverage in rubrics
Real-world rating hits boundary cases constantly; if your rubric doesn't address them, raters either skip work or guess, both poisoning the training data.
For example: “During a sports domain project, raters kept escalating questions about how to score plays where the rules seemed ambiguous. How would you have identified this category beforehand and expanded the rubric?”
3.Cross-domain translation skill
Each domain has its own jargon and conventions; you need to extract the core rule and restate it in language anyone outside that field can follow.
For example: “A retail specification mentions inventory reconciliation in accounting terminology. What steps would you take to confirm the raters understood it before deployment?”
4.Stakeholder alignment
Guidelines live or die on buy-in from both the research team writing them and the subject-matter experts checking them; misaligned expectations cause rework.
For example: “You discovered that the legal expert and the research program lead had opposite interpretations of a key clause. Describe how you would have surfaced and resolved that early.”
A task you may get
Given a draft rubric with at least two internal contradictions and one coverage gap, identify all three, propose fixes, and explain your reasoning in 500-750 words.
How to prepare
- Locate and review one published RLHF dataset or guideline paper to understand how large-scale rater instructions are structured
- Bring 2-3 concrete examples of ambiguity you have personally resolved, with before-and-after documentation
- Read about at least one GenAI or RLHF project publicly described to understand the workflow you'll be improving
The facts
- Pay
- $45–65/hr
- Hours
- Part time, 35 hours a week
- Where
- Remote · United States
- Open to
- USA
- Field
- Language and Audio
- Posted
- 7/10/2026
- Places left
- 10
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.