Training Turk

Bilingual German Generalist Expert — AI Safety

$48–52/hr · Mercor · Part time, 7 hours a week

You evaluate and strengthen AI model safety in German by writing test prompts, classifying content, and flagging adversarial patterns using structured guidelines and cultural judgment.

What you would do

  • Write sophisticated test prompts in German that explore model responses across sensitive and complex topics
  • Apply detailed classification systems to categorize prompts and responses according to structured safety guidelines
  • Identify adversarial phrasings, incremental boundary-pushing, and circumvention attempts in conversations
  • Flag content involving dual-use or sensitive information with documented reasoning
  • Document all classifications and concerns with clear explanations tied to safety principles

Who they want

  • Native or near-native fluency in German with strong written business-level English capability
  • Bachelor's degree completed or in progress
  • Strong written reasoning skills and high attention to detail in classification and annotation tasks
  • Sound judgment about sensitive and dual-use information handling
  • Ability to work independently on structured annotation tasks with consistency

Main skills

German fluencyPrompt engineeringContent classification

What the interview asks about

  1. 1.Prompt sophistication and cultural appropriateness

    Test prompts must be realistic and nuanced to reveal genuine model weaknesses; culturally tone-deaf or awkwardly phrased prompts fail to expose actual safety issues.

    For example: “Write two test prompts in German that probe how a model handles advice on a sensitive personal topic. One should use direct language, the other should employ indirection and cultural context. Explain why both variations matter for safety testing.”

  2. 2.Classification consistency under ambiguity

    Real content often sits between categories; maintaining consistent judgment across hundreds of cases requires disciplined application of guidelines, not intuition.

    For example: “Your guidelines define five levels of escalation in adversarial phrasings. You encounter a prompt that contains elements of level 2 and level 3. Walk me through how you'd apply the taxonomy and document your reasoning for a borderline call.”

  3. 3.Adversarial pattern detection

    Users often nudge boundaries incrementally; spotting escalation sequences and indirect circumvention attempts is core to evaluating model robustness.

    For example: “You're reviewing a conversation sequence: first prompt is innocent, second adds a constraint, third asks a related follow-up, fourth requests the same thing in different words. How would you flag the progression, and what does it reveal?”

  4. 4.Dual-use and sensitive content judgment

    Not all sensitive content requires flagging; sound judgment distinguishes between information that's sensitive but appropriate to discuss and material that creates genuine dual-use risk.

    For example: “A prompt asks for general information about a controlled process. Should you flag it? What specific factors would change your judgment, and how would you document your reasoning?”

  5. 5.Written documentation clarity

    Your written explanations become training data for researchers; unclear or vague reasoning defeats the purpose and may mislead model improvement efforts.

    For example: “Flag a prompt you judge adversarial and write a one-paragraph explanation that names the specific phrase or pattern, ties it to your classification guidelines, and explains why it matters for model safety.”

A task you may get

You receive a batch of 5 prompts in German and produce structured annotations for each, classifying them according to provided guidelines, flagging any adversarial elements, assessing dual-use content, and documenting your reasoning for all decisions.

How to prepare

  • Review provided classification guidelines thoroughly and study examples of each category to build intuition for consistent application
  • Practice writing test prompts in German on sensitive topics that are realistic, nuanced, and designed to probe model behavior
  • Research what dual-use information looks like and study examples of content that requires heightened caution
  • Read published work on adversarial prompting and jailbreak patterns to recognize escalation tactics and boundary-pushing sequences

The facts

Pay
$48–52/hr
Hours
Part time, 7 hours a week
Where
Remote · Remote — Western Europe preferred
Posted
9/4/2026

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.