$50–75/hr · Mercor · Part time, 40 hours a week
You use deep consumer or lifestyle expertise to evaluate and grade AI model recommendations, ensuring they are practical, accurate, safe, and culturally sound.
What you would do
- Compose realistic consumer scenarios with specific constraints, context, and a reference answer that represents sound judgment
- Review AI-generated recommendations and assess them on accuracy, practicality, budget alignment, and cultural appropriateness
- Detect and flag recommendations that sound confident but overlook real-world availability, pricing, seasonality, or social norms
- Write clear explanations of your evaluation reasoning to support the training process for AI models in your category
Who they want
- Substantive hands-on or expert-level background in a consumer category (sports, food, travel, fashion, entertainment, home technology, or related field)
- Clear written communication, since most work involves explaining your judgment and reasoning
- Ease with underspecified tasks and capability to surface gaps or inconsistencies in instructions
- Awareness of current market trends, pricing, availability, and regional or cultural variations in consumer behavior
- Part-time flexible availability; projects vary in scope and hours on short notice
Main skills
What the interview asks about
1.Evaluating practical real-world viability
Models often recommend products or approaches that sound good but are unavailable, out of budget, or seasonally irrelevant; catching these errors prevents consumers from wasting time or money.
For example: “An AI recommends a specific $300 hiking boot for someone with a $100 budget and a question asked in November for a February trip. Walk through your evaluation: is the recommendation wrong, or is there salvageable insight despite the flawed specifics?”
2.Cultural and regional appropriateness
Generic or US-centric advice fails for consumers in different cultural contexts; experts catch these blind spots before models embarrass users.
For example: “An AI recommends casual workplace accessories appropriate in America but not the consumer's formal business culture. How do you judge whether the core principle is salvageable or the whole recommendation misses the mark?”
3.Confidence and overreach detection
Models frequently assert false details with unwarranted certainty; identifying this pattern prevents users from acting on bad advice they trust too much.
For example: “An AI confidently states a discontinued product is available and recommends it as the optimal choice. How would you evaluate and document this error so the model learns to distinguish between products it knows well versus those it's guessing about?”
4.Writing evaluable scenarios
Vague consumer questions yield vague AI answers that are hard to grade; clear scenarios with specified constraints generate answers that can be fairly assessed.
For example: “Create a food scenario: dinner for 4, vegetarian, under 30 minutes, from typical grocery ingredients. Draft a reference answer and explain what details make a strong grading rubric.”
5.Harm and safety awareness
Recommendations about fitness, cooking, or travel can cause real problems if they overlook allergies, fitness levels, health conditions, or safety basics.
For example: “An AI recommends a high-intensity interval workout to someone without specifying they need medical clearance if they're recovering from injury. How would you flag this gap in your evaluation so the model becomes more cautious about fitness advice?”
A task you may get
Write a consumer scenario in your area of expertise (budget, constraints, context), draft a reference answer grading rubric, then evaluate a sample AI recommendation against it, explaining your judgments in 3-4 sentences.
How to prepare
- Document your background and 3-5 specific examples of consumer decisions or recommendations you've made in your category, noting the constraints or trade-offs that mattered
- Track recent category changes (new products, pricing, availability, trends) and practice assessing whether an AI would catch them
- Draft 2-3 realistic consumer questions with constraints (budget, time, availability, preferences) and think about what a good answer would need to address
- Review how you'd spot overconfident errors in recommendations and practice writing clear, concise evaluations that explain why an answer fails or succeeds
The facts
- Pay
- $50–75/hr
- Hours
- Part time, 40 hours a week
- Where
- Remote
- Field
- Miscellaneous
- Role type
- Talent network
- Posted
- 9/15/2026
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.