Training Turk

Personalized Life Assistant Expert

$50–200/hr · Mercor · Hourly, 40 hours a week

A Personalized Life Assistant Expert evaluates LLM outputs to improve how AI systems handle real-world personal tasks.

What you would do

  • Rate LLM responses on effectiveness for personal life tasks including planning, research, and decision-making scenarios
  • Judge whether AI output demonstrates context-awareness, personalization, realism, and safety for the user's situation
  • Evaluate multi-step AI-assisted workflows to identify gaps in logic, missing constraints, or poor tradeoff analysis
  • Document detailed written feedback explaining what makes an AI response succeed or fail in practice
  • Help train AI systems to better understand real-world personal workflows and user success criteria

Who they want

  • Heavy personal experience using LLM products like ChatGPT, Claude, Gemini, Perplexity, Cursor, or similar tools
  • Demonstrated ability to use AI for complex personal tasks spanning planning, research, decision-making, and workflow optimization
  • Strong written communication skills with excellent attention to detail in explaining AI performance
  • Ability to explain what distinguishes effective AI responses from weak, incomplete, unsafe, or implausible ones
  • Availability for 20-40 hours per week of evaluation work

What the interview asks about

  1. 1.Personalization and context awareness

    Generic advice fails users; evaluators must spot whether AI actually understood personal constraints, preferences, and success metrics in specific situations.

    For example: “An AI suggests a career pivot strategy for someone with 10 years in finance, two young kids, and a mortgage. How would you evaluate whether the advice accounts for that person's constraints, or whether it's generic guidance that ignores their reality?”

  2. 2.Multi-step reasoning in personal workflows

    Personal AI tasks often span multiple steps with dependencies; evaluators judge whether AI chains reasoning correctly or breaks down partway through.

    For example: “Someone asks AI to help plan a kitchen renovation while managing a full-time job. What would you assess to determine whether the AI response actually creates a realistic, actionable plan or just lists generic considerations?”

  3. 3.Identifying unsafe or unrealistic outputs

    AI can produce plausible-sounding advice that is actually dangerous or impossible to execute; evaluators must catch this before it influences real decisions.

    For example: “An AI recommends a specific health protocol to improve energy levels. What checks would you perform to determine whether the recommendation is evidence-based, realistic for a busy person to follow, or potentially harmful?”

  4. 4.Assessing AI tool strengths and limitations

    Effective evaluators know what each AI tool excels at and where it typically fails, so they judge output quality in context of the tool's actual capabilities.

    For example: “You're evaluating an AI's response to a complex financial planning question involving taxes, retirement accounts, and investment strategy. How would you account for the AI's limitations in giving personalized financial advice?”

  5. 5.Detail-oriented written judgment

    Evaluation feedback must be precise and actionable for AI trainers; vague or incomplete assessments don't drive improvements.

    For example: “An AI response for meal planning seems reasonable but generic. What specific observations would you include in your assessment to help trainers understand where personalization failed?”

  6. 6.Evaluating decision-making support

    Personal AI assistants should help users make better decisions by surfacing tradeoffs and alternatives, not just give prescriptive answers.

    For example: “Someone asks AI for help deciding between two job offers. How would you judge whether the AI response actually helped them think through the decision systematically, or just narrowed it down too early?”

A task you may get

Review an AI response to a multi-step personal task (e.g., moving to a new city, planning education and career). Identify strengths, gaps in personalization, unrealistic assumptions, and missing context. Write a detailed evaluation explaining your judgments.

How to prepare

  • Use at least 3 different LLM tools for a week in real personal workflows to develop intuition about their strengths and blind spots
  • Study examples of what constitutes good vs bad personalized advice across domains like health, career, and learning
  • Practice writing detailed, specific feedback on AI outputs that explains exactly what is missing or incorrect, not just ratings
  • Research how AI systems fail at context-awareness and personalization in real-world user studies

The facts

Pay
$50–200/hr
Hours
Hourly, 40 hours a week
Where
Remote
Open to
USA, USA
Field
Life, Physical, and Social Science
Project name
Dorado
Posted
6/13/2026
Places left
100

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.