LLM Red-Teamer for AI Model Evaluation
$40–$65/hr
About this role
Design and execute adversarial, multi-turn prompts and task scenarios to probe and improve frontier language models. Your work will produce evaluation packages and validated assessments that reveal model strengths and failure modes, and directly inform model training and behavior improvements.
Key Responsibilities- Develop complex, adversarial multi-turn conversations and task-based scenarios that follow detailed project specifications.
- Author clear, precise evaluation rubrics to assess model responses against defined behavioral targets.
- Iteratively test scenarios against state-of-the-art LLMs, increasing difficulty and nuance until quality thresholds are reached.
- Deliver comprehensive task packages, including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
- Validate LLM outputs, documenting model strengths and specific failure modes relative to the project specification.
- Maintain calibration with team leads and quality control contacts as requirements evolve.
- Work independently to produce high-quality deliverables at a steady, consistent pace.
- Required skills: adversarial prompt construction, precision in written English, rubric design, and iteration stamina.
- No prior experience in AI is required; domain expertise and real-world subject matter knowledge are valued.
- Exceptional written English, including clarity, precision, and strong organization.
- Ability to interpret and execute complex specifications with minimal oversight, strong critical thinking, and meticulous attention to detail.
- Experience designing evaluation items or rubrics is advantageous.
- Backgrounds that map well to this role include research, editorial work, technical writing, analysis, or quality assurance.
- Familiarity with large language models and common failure patterns is a plus.
- Role type: Contractor, remote.
- Compensation is output-based; experts are paid per task that meets the project specifications.
- The time required to complete work will vary by expert experience and workflow.
- Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
- We typically fill roles within 48 hours and look for experts who can begin quickly.
- If selected, you should be ready to complete onboarding and start your first tasks within 24 to 48 hours after onboarding is finished.
- Listed rate: $40 to $65 per hour.
- Actual payment is tied to completed tasks that meet specifications, and effective hourly earnings will depend on task throughput and speed.
Applications are reviewed rapidly. If selected, you will complete a short onboarding process, after which you are expected to begin contributing to tasks within 24 to 48 hours. Candidates who can start immediately are preferred.
EligibilityNo specific work-authorization or citizenship restrictions were specified in the source. Applicants should be able to work as independent contractors and manage remote, self-directed assignments.