Training Turk

PDF Annotation & Transcription Experts – Bengali

$10–11/hr · Mercor · Hourly, 20 hours a week

You annotate Bengali PDF pages with structural markup and produce faithful character-level transcriptions to train AI document-understanding systems.

What you would do

  • Identify and bound every meaningful region on the page (titles, sections, paragraphs, lists, tables, figures)
  • Classify each region consistently using a defined component taxonomy
  • Document reading order and relationships between regions
  • Transcribe all text exactly as it appears in native Bengali script and diacritics
  • Flag regions with handwriting or legibility issues and record page metadata

Who they want

  • Native Bengali speaker with complete command of Bengali script and diacritics
  • Background in annotation, transcription, translation, localization, or proofreading
  • Character-level accuracy orientation; precision valued over speed
  • Comfort with unfamiliar layouts including multi-column pages, handwritten forms, and mixed-script documents
  • Familiarity with document production, OCR, or regional-language AI data work preferred

What the interview asks about

  1. 1.Bengali script and orthography mastery

    AI models learn from transcription accuracy; this tests whether your Bengali command includes diacritics, conjunct forms, and numeral variants.

    For example: “A Bengali page contains a word with a complex conjunct form and multiple diacritics. You're unsure whether you're reading it correctly. How would you verify the correct transcription?”

  2. 2.Complex page layout navigation

    Real Bengali documents have multi-column layouts, sidebars, figures with captions, and mixed scripts; this tests whether you can handle that complexity.

    For example: “A newspaper page has three columns of text, a sidebar with a different font, a photograph with a caption in Bengali, and an embedded advertisement. How would you annotate the reading order and region boundaries?”

  3. 3.Taxonomy consistency over hundreds of pages

    Dataset quality requires consistent classification; this tests whether you apply the same rules even when documents vary widely.

    For example: “You've annotated 200 pages. You notice that what you called a 'caption' on page 1 could also be classified as a 'note' or 'label.' How would you decide and how would you handle earlier pages?”

  4. 4.Handwriting recognition and legibility judgment

    Training data must accurately reflect what can and cannot be read; this tests whether you distinguish handwriting confidence from evasion.

    For example: “A page has handwritten Bengali text in the margin. The handwriting is legible but ambiguous in a few characters. Do you transcribe your best guess, flag it as unclear, or skip it?”

  5. 5.Quality review and error identification

    Every annotation is peer-reviewed; this evaluates whether you can spot errors in others' work systematically.

    For example: “You're reviewing a colleague's annotations of a table-heavy page. You notice they've classified one column header as a separate region, but you think it belongs to the table. How would you document that concern?”

A task you may get

Annotate a 1-2 page Bengali document sample, identifying regions, applying consistent component classification, and transcribing all text accurately in Bengali script.

How to prepare

  • Review a Bengali document style guide or annotation taxonomy to understand classification principles
  • Collect 3-4 real Bengali documents in various formats (newspaper, textbook, form) to practice on
  • Prepare a checklist of common Bengali orthographic elements to verify your accuracy

The facts

Pay
$10–11/hr
Hours
Hourly, 20 hours a week
Where
Remote
Open to
IND
Field
Miscellaneous
Project name
Neon
Posted
9/2/2026
Places left
10

We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.