LLM Red-Teamer for AI Model Evaluation
$40–$65/hr
About this role
Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn prompts, designing evaluation rubrics, and delivering high-quality task packages that expose model strengths and failure modes. This contractor role focuses on improving how next-generation AI systems learn, reason, and behave, and does not require prior AI experience; strong domain knowledge and writing ability are the primary requirements.
Key Responsibilities- Develop complex, adversarial multi-turn conversations and task-based scenarios that follow detailed project specifications.
- Author clear, precise evaluation rubrics to assess model responses against defined behavioral targets.
- Iteratively test conversations and tasks on frontier LLMs, increasing difficulty and nuance until the desired quality threshold is met.
- Deliver complete task packages including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
- Validate and document LLM outputs, identifying model strengths, failure modes, and deviations from the project specification.
- Maintain calibration with team leads and quality control contacts as project requirements evolve.
- Work independently to produce steady, consistent, high-quality deliverables.
- Required skills: adversarial prompt construction, precision in written English, rubric design, and iteration stamina.
- Exceptional written English ability, with clarity, precision, and strong structural organization.
- No prior AI experience required; domain expertise is valued.
- Preferred: prior experience in human data environments such as RLHF, SFT, evaluations, annotation, or prompt engineering.
- Preferred: deep familiarity with large language models and common failure patterns.
- Proven ability to interpret and execute complex specifications with minimal oversight, strong critical thinking, and high attention to detail.
- Experience designing evaluation items or rubrics, or backgrounds in writing-intensive or analysis-centric fields such as research, editorial, technical writing, or quality assurance are advantages.
- Role type: Contractor, remote.
- Compensation is output-based, with experts paid per task that meets project specifications.
- Pay reference: $40 to $65 per hour.
- The time required to complete tasks will vary by expert experience and workflow.
- Minimum submission requirements apply, experts must submit a minimum of tasks per week.
Pay range reference: $40 to $65 per hour. Actual payments are task-based, awarded for deliverables that satisfy the project specifications. Time-to-complete will vary by individual. Minimum submission requirements apply.
Eligibility and Application Process- This is a remote independent contractor engagement.
- We typically fill roles within 48 hours and seek experts who can begin quickly.
- If selected, you should be prepared to start your first tasks within 24 to 48 hours after completing onboarding.
- Onboarding must be completed before you can begin tasks, and selection is required prior to onboarding.