$31/hr · Mercor · Hourly, 20 hours a week
You annotate and transcribe Japanese-language PDF pages, mapping structure and transcribing text with character-level accuracy to build training data for document-understanding AI.
What you would do
- Analyze PDF pages and identify every meaningful content region: titles, sections, paragraphs, lists, tables, figures, diagrams, captions, and formulas
- Assign each region a component type and reading-order index, capturing relationships to parent components
- Transcribe all text faithfully in original Japanese script, including all character types, variant forms, and handwritten passages
- Record page metadata: language, document type, source, dimensions, and flags for special features like tables and handwriting
- Review colleague work systematically, verifying region identification, transcription exactness, and metadata accuracy
Who they want
- Native Japanese speaker with complete fluency in kanji, hiragana, katakana, including furigana and character variants
- Work history in linguistic or document-focused roles like media, translation services, proofreading, or text digitization
- Commitment to character-level accuracy and systematic taxonomy application across extended annotation tasks
- Ease handling diverse and challenging page structures: vertical scripts, side-by-side columns, test documents, and written annotations
- Experience with AI training data, annotation, OCR correction, transcription, localization, or related quality work
Main skills
What the interview asks about
1.Character accuracy and variant handling
A single wrong character breaks training data quality; you must recognize variant forms, encoding edge cases, and handwritten shapes accurately without cutting corners.
For example: “A newspaper page contains a kanji character that appears in multiple variant forms across different regions, plus furigana readings above some instances and not others. How would you ensure consistent representation in transcription?”
2.Reading order determination
Correct reading order determines how AI learns to parse page logic; vertical text, multi-column layouts, and mixed orientations require systematic judgment, not guessing.
For example: “You're annotating a three-column newspaper page with a headline spanning all columns and text flowing vertically within each column. Map out your reading-order sequence and explain the logic behind it.”
3.Component type classification
Misclassifying regions teaches the model wrong structure; you must apply taxonomy consistently rather than improvising per document or layout type.
For example: “You encounter a page with numbered lines, side annotations, highlighted quotes, and traditional tables. Walk through how you'd classify each element type and assign reading order when the structure is unconventional.”
4.Handwriting transcription and flagging
Handwritten regions add training value but require accuracy and honesty; illegible sections must be flagged rather than guessed or skipped.
For example: “A form contains a handwritten signature, a handwritten date in cursive, and handwritten text in cramped spacing that's partially legible. How would you handle each, and what would you flag?”
5.Unsuitable page identification
Catching unsuitable pages early saves work; you should screen pages for language, legibility, and personal information before investing thirty minutes in annotation.
For example: “You open a page that appears to be in Japanese but contains a significant block of text in Chinese characters that dominates the page. Is this suitable for the Japanese dataset? What additional checks might you perform?”
A task you may get
You receive a complex Japanese PDF with multi-column layout, vertical text, handwriting, and tables. Produce complete annotations including region identification, component typing, reading order, parent relationships, transcription, and metadata.
How to prepare
- Study the full annotation taxonomy: learn every component type, how reading order works in vertical and multi-column layouts, and when to flag special features
- Practice transcribing Japanese documents with mixed script, handwriting, and variant characters; build accuracy and speed on your own time before formal work
- Research Unicode normalization and Japanese encoding standards so you understand variant forms and encoding choices
- Review examples of annotated pages in your dataset to calibrate expectations for region boundaries, reading order, and transcription standards
The facts
- Pay
- $31/hr
- Hours
- Hourly, 20 hours a week
- Where
- Remote
- Open to
- JPN, USA, CAN, GBR
- Field
- Language and Audio
- Project name
- Neon
- Posted
- 9/2/2026
- Places left
- 3
We wrote this page from the public Mercor listing. It may be out of date, so read the full posting before you apply.