Key details
- Role type
- Contract
- Compensation
- $33.58/hr
- Work arrangement
- Remote
- Category
- language
- Confirmed requirements
- 4
About this role
Role Overview
Help build high-quality training data for document AI by creating detailed structural annotations and exact Japanese transcriptions from real, publicly available PDF pages. You will work with complex material such as handwriting, dense multi-column layouts, vertical text, tables, diagrams, and mixed-script documents drawn from newspapers, textbooks, examinations, flyers, forms, manuals, menus, brochures, notices, and worksheets.
All annotation and transcription is completed in Japanese and is human-authored, including component identification, component classification, reading order, and transcription.
Key Responsibilities
- Find publicly accessible Japanese PDFs in assigned document categories that include at least one multimodal element, such as an image, table, diagram, or handwriting, and document the source.
- Identify, bound, classify, and assign reading-order indexes to every meaningful page region, including titles, headings, paragraphs, lists, tables, figures, diagrams, captions, formulas, questions, and answer fields.
- Link page regions to their related figures or tables using parent component identifiers.
- Transcribe text exactly as shown, including kanji, hiragana, katakana, furigana, and handwritten content, and flag illegible source regions.
- Capture page metadata, including language, document type, source, page dimensions, and indicators for tables, formulas, and handwriting.
- Review a colleague’s work end to end. Experienced annotators may take on second-expert review responsibilities.
Qualifications
- Native Japanese fluency with full command of kanji, hiragana, katakana, furigana, and variant character forms.
- Professional experience in interpretation, journalism, transcription, translation, editorial work, or similar document-intensive work.
- Exceptional character-level accuracy and a quality-first approach.
- A systematic approach to applying a consistent taxonomy across hundreds of pages.
- Comfort working with unfamiliar document formats, including vertical text, multi-column newspapers, examination papers, and handwritten forms.
Preferred Experience
- AI training-data annotation, labeling, grading, or bilingual evaluation.
- MTPE, subtitling, bilingual quality assurance, OCR correction, or post-editing.
- Japanese-language typesetting, copy-editing, proofreading, or digitization.
- Unicode normalization, Japanese input methods, full-width and half-width forms, or kanji variant handling.
What Success Looks Like
- Every meaningful page region is accurately captured, bounded, and typed.
- Reading order reflects how the document is actually read, including vertical and multi-column layouts.
- Japanese transcriptions match the source character for character, rather than using romaji.
- Tasks pass second-expert review on the first submission.
- Source documents add meaningful layout diversity to the corpus.
Work Terms
- Remote, hourly engagement.
- Work is performed in Japanese.
Compensation
$33.58 per hour.
What to prepare before applying
- Native fluency in Japanese with full command of kanji, hiragana, katakana, furigana, and variant character forms.
- Professional experience in interpretation, journalism, transcription, translation, editorial work, or comparable document-intensive work.
- Character-level accuracy in Japanese transcription.
- Ability to apply a document-component taxonomy consistently across many pages.
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.