Key details
- Role type
- Contract
- Compensation
- $12.68/hr
- Work arrangement
- Remote — US, Canada or India
- Category
- language
- Confirmed requirements
- 4
About this role
Role Overview
Help create human-authored training data that improves document AI for Gujarati script. You will map the full structure of real, publicly available Gujarati PDF pages and produce faithful Gujarati-script transcriptions for every text region, including challenging content such as handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages.
Documents may include newspapers, textbooks, examinations, flyers, forms, manuals, menus, brochures, notices, and worksheets. Accuracy is the primary measure of quality, with handling time also tracked.
Key Responsibilities
- Find publicly accessible Gujarati PDFs in assigned document categories that include at least one image, table, diagram, or handwritten element, and record the source.
- Identify, bound, and classify every meaningful page region, including titles, headings, paragraphs, lists, tables, figures, diagrams, captions, formulas, questions, and answer fields.
- Assign reading-order indices and link related regions to their parent figure or table.
- Transcribe all text exactly as shown in Gujarati script, including handwritten text, and flag content that is not legible.
- Record page metadata, including language, document type, source, page dimensions, and indicators for tables, formulas, and handwriting.
- Review another contributor''s work end to end when assigned as an experienced annotator.
Qualifications
- Native fluency in Gujarati, with complete command of Gujarati script, orthography, diacritics, and conjunct forms.
- All annotation and transcription work is performed in Gujarati.
- Experience with document-focused work such as annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
- Character-level precision, including accurate diacritics, is essential.
- Ability to apply a consistent taxonomy across many pages and work confidently with unfamiliar layouts, including multi-column newspapers, exam papers, and handwritten forms.
Preferred Experience
- Regional-language AI data work, including annotation, labeling, grading, or bilingual evaluation for training datasets.
- MTPE, subtitling, bilingual QA, OCR correction, or post-editing.
- Gujarati document production, typesetting, copy-editing, proofreading, or digitization.
- Unicode normalization, Gujarati input methods, and numeral and character-form accuracy.
Work Terms
- Remote hourly engagement.
- Open only to contributors based in India, the United States, Canada, or Western Europe.
Compensation
- $12.68 per hour.
Quality Expectations
- Capture, correctly bound, and correctly classify every meaningful region on each page.
- Reflect the true reading order, including across columns.
- Deliver character-for-character Gujarati-script transcriptions, without transliteration.
- Produce work that passes second-expert review on the first submission.
- Contribute documents with varied layouts rather than repeating existing templates.
What to prepare before applying
- Be based in India, the United States, Canada, or Western Europe.
- Native fluency in Gujarati, including full command of Gujarati script and orthography.
- Full command of Gujarati diacritics and conjunct forms.
- Experience working with documents, such as annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.