Training Turk

Bilingual Czech Generalist Expert — AI Safety

$43–47/hr · Mercor · Part time, 7 hours a week

A Czech language expert who designs challenging test scenarios in Czech and rates whether AI models handle sensitive topics safely, without requiring machine learning background.

What you would do

  • Author sophisticated prompts in Czech designed to test model behavior across sensitive social and harmful-content domains
  • Rate model outputs using structured safety rubrics, identifying violations and risky reasoning patterns
  • Detect and document when users employ escalation tactics or attempt to bypass safety measures through subtle reformulation
  • Document your analysis and reasoning with clarity so that engineering teams understand why particular responses require revision

Who they want

  • Bachelor's degree (completed or actively pursuing)
  • Native or near-native fluency in Czech combined with professional-standard English writing ability
  • Strong capacity for precise written communication and meticulous attention to detail
  • Mature judgment regarding sensitive topics and the potential consequences of releasing particular information

Main skills

Czech language expertiseSensitive topic evaluationAdversarial prompt detection

What the interview asks about

  1. 1.Constructing adversarial Czech prompts

    Test prompt quality directly determines whether model safety measures are genuine or merely surface-level; weak prompts allow unsafe models to appear compliant, corrupting the evaluation.

    For example: “Design three progressively more challenging Czech prompts about a sensitive political topic, where each iteration subtly increases pressure on the model to violate guidelines. Explain how your framing shifts without becoming obviously hostile.”

  2. 2.Recognizing cultural context in escalation

    A manipulation sequence that works in one language or culture may fail in another; Czech-specific social dynamics and communication norms affect how adversarial tactics succeed or fail.

    For example: “You observe a sequence where a user employs indirect appeals to Czech nationalism to encourage the model to relax content policies. Describe how the escalation works and why it might be more effective in Czech than in English.”

  3. 3.Distinguishing safety violations from edge cases

    Over-flagging creates excessive false positives that waste engineering time; under-flagging allows genuinely harmful outputs into production; your judgment must hit the narrow middle path.

    For example: “A model provides factual information about a topic classified as sensitive but refuses to draw harmful conclusions. Using the safety rubric, explain whether this qualifies as a violation or an appropriate boundary.”

  4. 4.Consistency across fatigue and variation

    After hundreds of evaluations, reviewers naturally drift in stringency; undetected drift corrupts the training signal, so you must recognize and counteract your own shifting standards mid-session.

    For example: “You're 50 prompts into a 200-prompt batch and notice you're flagging fewer cases. What evidence would make you recognize this drift, and what corrective steps would you implement?”

  5. 5.Translating safety concerns to engineers

    Engineers cannot redesign models based on vague complaints; your documentation must articulate exactly why a response failed, what specific guideline applies, and what alternatives would succeed.

    For example: “A model's response is technically accurate but uses social proof and peer pressure to justify a policy violation. How would you describe this to engineers so they understand the flaw and can fix it?”

A task you may get

Review a series of 15 Czech prompts with corresponding model responses, classify each using the provided safety rubric, and write detailed explanations for your top-three highest-risk and top-three most-concerning edge cases.

How to prepare

  • Practice documenting safety judgments in clear, evidence-based language so your reasoning is easy for others to verify and build upon
  • Study examples of content moderation frameworks and escalation tactics to build pattern recognition for adversarial user behavior
  • Review recent news and cultural discussions in Czech media to refresh your understanding of sensitive contemporary topics

The facts

Pay
$43–47/hr
Hours
Part time, 7 hours a week
Where
Remote · Remote — Eastern Europe preferred
Posted
9/4/2026

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.