About this role
Design and test challenging, multi-turn scenarios that help improve how frontier language models learn, reason, and respond. This remote contract role focuses on producing rigorous evaluations and actionable feedback through high-quality adversarial prompts, task design, and model analysis.
Key Responsibilities- Create complex adversarial conversations and task-based scenarios that follow detailed project specifications.
- Write clear, precise evaluation rubrics that assess model responses against defined behavioral targets.
- Test tasks against frontier language models, iterating to increase difficulty and nuance until quality requirements are met.
- Deliver complete task packages, including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
- Validate model outputs and document strengths and failure modes against project requirements.
- Stay aligned with team leads and quality-control contacts as requirements evolve.
- Work independently and maintain a steady pace of high-quality deliverables that support next-generation AI training.
- Strong adversarial prompt-construction skills.
- Exceptional written English, with clarity, precision, and strong organization.
- Experience designing evaluation items or rubrics, or the ability to do so effectively.
- Comfort iterating through tests and refinements with sustained attention to detail.
- Ability to interpret complex specifications and work autonomously with minimal oversight.
- Critical-thinking skills and meticulous attention to detail.
- Experience with RLHF, SFT, evaluations, annotation, prompt engineering, or other AI human-data work is preferred.
- Familiarity with large language models and common model failure patterns is preferred.
- Experience in research, editorial work, technical writing, quality assurance, or other writing-intensive or analysis-focused fields is a plus.
- No prior AI experience is required; relevant domain expertise is valued.
- Remote, independent contractor engagement.
- Work is output-based, and time required may vary based on experience and workflow.
- Minimum submission requirements apply, including a minimum number of tasks submitted each week.
- Selected candidates should be ready to begin their first tasks within 24 to 48 hours after completing the required setup process.
- Advertised compensation range: $40 to $65 per hour.
- Payment is based on completed tasks that meet project specifications.