Key details
- Role type
- Contract
- Compensation
- $12.68/hr
- Work arrangement
- Remote
- Category
- language
- Confirmed requirements
- 4
About this role
Role Overview
Help build high-quality training data for document AI by creating complete structural annotations and faithful Gujarati-script transcriptions from real, publicly available PDF pages. You will work with challenging materials such as handwriting, multi-column layouts, tables, diagrams, and mixed-script pages drawn from newspapers, textbooks, examinations, flyers, forms, manuals, menus, brochures, notices, and worksheets.
All annotation and transcription work is performed in Gujarati. Accuracy is the primary quality measure, with handling time also tracked.
Key Responsibilities
- Find publicly accessible Gujarati PDFs within assigned document types that include at least one image, table, diagram, or handwritten element, and document the source.
- Identify, bound, classify, and assign reading-order indexes to every meaningful page region, including titles, headings, paragraphs, lists, tables, figures, diagrams, captions, formulas, questions, and answer fields.
- Connect regions to their related figures or tables using parent component identifiers.
- Transcribe all text exactly as shown in Gujarati script, including handwriting, and flag text that cannot be read clearly.
- Record page metadata, including language, document type, source, page dimensions, and whether tables, formulas, or handwriting are present.
- Review another Gujarati expert''s completed work end to end when assigned as an experienced annotator.
Qualifications
- Native Gujarati fluency, including full command of Gujarati script, orthography, diacritics, and conjunct forms.
- Experience with document annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
- Character-level precision and a consistent approach to applying a taxonomy across many pages.
- Comfort working with unfamiliar layouts, including multi-column newspapers, examination papers, and handwritten forms.
Preferred Experience
- Regional-language AI data annotation, labeling, grading, or bilingual evaluation.
- MTPE, subtitling, bilingual quality assurance, OCR correction, or post-editing.
- Gujarati document typesetting, copy-editing, proofreading, or digitization.
- Unicode normalization, Gujarati input methods, and accurate numeral and character forms.
What Success Looks Like
- Every meaningful page region is captured, correctly bounded, and correctly classified.
- Reading order matches how the document is read, including across columns.
- Gujarati transcriptions match the source character for character and are not transliterated.
- Work passes second-expert review on the first submission.
- Submitted documents add meaningful layout diversity to the dataset.
Work Terms
- Remote, hourly engagement.
Compensation
- $12.68 per hour.
Eligibility
- Gujarati fluency is required. Candidates must be able to read and write Gujarati script accurately for all annotation and transcription tasks.
What to prepare before applying
- Native fluency in Gujarati, including full command of Gujarati script and orthography.
- Full command of Gujarati script, including diacritics and conjunct forms.
- Experience working with documents, such as annotation, transcription, translation, localization, subtitling, proofreading, journalism, or regional-language data review.
- Character-level accuracy in Gujarati transcription, including accurate diacritics.
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.