Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

LLM Red-Teamer for AI Model Evaluation

$40–$65/hr

RemoteRemote micro1 is engaging LLM Red-Teamers to contribute to a high-impact customer project focused on the evaluation and improvement of frontier language models. In this role, you'll apply your expertiContracttechnology
Apply Now

About this role

Role Overview

Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn prompts, designing evaluation rubrics, and delivering high-quality task packages that expose model strengths and failure modes. This contractor role focuses on improving how next-generation AI systems learn, reason, and behave, and does not require prior AI experience; strong domain knowledge and writing ability are the primary requirements.

Key Responsibilities
  • Develop complex, adversarial multi-turn conversations and task-based scenarios that follow detailed project specifications.
  • Author clear, precise evaluation rubrics to assess model responses against defined behavioral targets.
  • Iteratively test conversations and tasks on frontier LLMs, increasing difficulty and nuance until the desired quality threshold is met.
  • Deliver complete task packages including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
  • Validate and document LLM outputs, identifying model strengths, failure modes, and deviations from the project specification.
  • Maintain calibration with team leads and quality control contacts as project requirements evolve.
  • Work independently to produce steady, consistent, high-quality deliverables.
Qualifications
  • Required skills: adversarial prompt construction, precision in written English, rubric design, and iteration stamina.
  • Exceptional written English ability, with clarity, precision, and strong structural organization.
  • No prior AI experience required; domain expertise is valued.
  • Preferred: prior experience in human data environments such as RLHF, SFT, evaluations, annotation, or prompt engineering.
  • Preferred: deep familiarity with large language models and common failure patterns.
  • Proven ability to interpret and execute complex specifications with minimal oversight, strong critical thinking, and high attention to detail.
  • Experience designing evaluation items or rubrics, or backgrounds in writing-intensive or analysis-centric fields such as research, editorial, technical writing, or quality assurance are advantages.
Work Terms
  • Role type: Contractor, remote.
  • Compensation is output-based, with experts paid per task that meets project specifications.
  • Pay reference: $40 to $65 per hour.
  • The time required to complete tasks will vary by expert experience and workflow.
  • Minimum submission requirements apply, experts must submit a minimum of tasks per week.
Compensation

Pay range reference: $40 to $65 per hour. Actual payments are task-based, awarded for deliverables that satisfy the project specifications. Time-to-complete will vary by individual. Minimum submission requirements apply.

Eligibility and Application Process
  • This is a remote independent contractor engagement.
  • We typically fill roles within 48 hours and seek experts who can begin quickly.
  • If selected, you should be prepared to start your first tasks within 24 to 48 hours after completing onboarding.
  • Onboarding must be completed before you can begin tasks, and selection is required prior to onboarding.

Related Jobs