AI Evaluation Analyst
$20–$30/hr
About this role
Contribute domain expertise to a client project that trains and evaluates next-generation large language models. In this role you will design multi-turn conversations and evaluation rubrics, test and refine prompts against frontier LLMs, and produce high-quality evaluation assets that shape how models learn, reason, and respond. No prior AI experience is required; strong written clarity, domain knowledge, and the ability to follow detailed specifications are what matter.
Key Responsibilities- Author detailed, task-based multi-turn conversations and associated rubrics aligned to project specifications.
- Test draft conversations and prompts against frontier large language models, iterate to meet required quality and difficulty levels.
- Deliver evaluation assets such as transcripts, target behaviors, binary rubrics, and supporting evidence for model outputs.
- Maintain strict fidelity to evolving project specifications while producing work at high throughput.
- Validate and calibrate outputs with team leads and quality control as guidelines change.
- Work independently and consistently to meet expected output rates for deliverable completion.
- Required skills: working knowledge of frontier LLM behavior, experience with data annotation, written English clarity and structure, ability to maintain spec fidelity at volume, and self-direction to follow detailed specifications.
- Preferred qualifications: native-level written English, prior experience in data annotation, RLHF, SFT, evaluation, or prompt engineering, demonstrated ability to interpret and apply highly detailed specifications without supervision, strong critical thinking and analytical skills in writing-heavy or analysis-heavy domains, and experience authoring evaluation items, rubrics, or performing deep analysis of technology outputs.
- Backgrounds in research, editorial, technical writing, or quality assurance are a plus.
- Role type: Contractor, remote.
- Compensation is output-based; experts are paid per task that meets the project specifications.
- Time to complete tasks varies by expert experience and workflow; minimum submission requirements apply.
- Experts must submit a minimum of tasks per week.
Pay rate range: $20.00 to $30.00 per hour (indicative). Actual payment is task-based, paid for tasks that meet the project specifications. The time required to complete work may vary depending on experience and workflow.
Eligibility and Application Process- Positions are remote and open to independent contractors who can comply with detailed specifications and work independently.
- We typically fill roles within 48 hours and seek experts who can begin quickly.
- If selected, you will complete onboarding and are expected to start your first tasks within 24 to 48 hours of finishing onboarding.