AI Evaluation Analyst
$20–$30/hr
About this role
Contribute domain expertise to train and evaluate frontier large language models by authoring multi-turn conversations, detailed rubrics, and evaluation assets that shape how next-generation AI systems learn, reason, and perform. This remote contractor role emphasizes written clarity, spec fidelity at scale, and independent work against evolving model behavior.
Key Responsibilities- Create detailed, task-based multi-turn conversations and corresponding rubrics that align precisely with project specifications.
- Test drafts against frontier LLMs, iterate to meet quality and difficulty requirements, and refine examples for consistency and challenge.
- Deliver evaluation materials including transcripts, identified target behaviors, binary rubrics, and supporting evidence for judgments.
- Maintain strict fidelity to evolving project specifications while producing work at the expected throughput.
- Validate and calibrate outputs with team leads and quality control as guidelines change.
- Work independently to meet expected output rates and deliverables.
- Required skills: working knowledge of frontier LLM behavior, experience with data annotation, strong written English clarity and structure, ability to maintain spec fidelity at volume, and self-direction to follow a detailed specification.
- Preferred: native-level written English; prior experience in data annotation, RLHF, SFT, evaluation, or prompt engineering; familiarity with common model failure patterns; experience authoring evaluation items, rubrics, or conducting deep analysis of technology outputs.
- Additional strengths: strong critical thinking, analytical ability in writing-heavy or analysis-heavy domains, and backgrounds in research, editorial, technical writing, or quality assurance.
- No prior AI job experience required, domain knowledge and subject-matter expertise are valued.
- Role type: Contractor, remote.
- Compensation is output-based, experts are paid per task that meets the project specifications.
- Time to complete tasks varies by experience and workflow.
- Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
Base compensation range: 20 - 30 hourly, paid per approved task that meets project specifications.
Application and Start Timeline- Roles are typically filled within 48 hours.
- If selected and onboarded, you will be expected to start your first tasks within 24 to 48 hours of completing onboarding.
micro1 is an AI data lab that converts subject-matter expertise into high-quality training data and evaluations to improve AI systems. Experts contribute across domains such as finance, healthcare, and engineering to build the human intelligence layer for frontier AI.