Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

LLM Red Team Specialist for AI Model Evaluation

$60–$90/hr

Remote — US onlyUnited StatesContracttechnology
Apply Now

About this role

Role Overview

Identify subtle failure modes and edge cases in frontier large language models by designing, executing, and analyzing complex multi-step probes. You will operate in a red-teaming setup to find situations where models appear competent but produce incorrect, unsafe, or misleading outputs, and convert those discoveries into reproducible benchmark tasks that improve model evaluation.

Key Responsibilities
  • Probe models, exploring behavior on coding, machine learning, and analytical tasks to surface quiet failures and hidden vulnerabilities.
  • Design challenges that capture discovered weaknesses, ensuring tasks are challenging for models while remaining fair to grade.
  • Document findings clearly, with evidence, reproduction steps, and analysis that other researchers can follow.
  • Collaborate with task authors and researchers to close loopholes, remove shortcuts, and improve grading robustness.
  • Share insights across the team to iteratively strengthen the benchmark and evaluation process.
Qualifications
  • MSc or PhD in a STEM field, or equivalent practical experience in a research-intensive role involving data analysis and coding.
  • At least 1 year of experience in research, research engineering, security, adversarial testing, or AI evaluation roles.
  • Proven ability to identify vulnerabilities, edge cases, or failure modes in LLMs or other ML systems through red teaming, adversarial testing, security research, or rigorous evaluation.
  • Working proficiency in Python and Git, with the ability to write scripts for probes and analyses.
  • Strong familiarity with LLM capabilities, limitations, and common evaluation techniques.
  • Preferred: prior experience in AI training, model evaluation, or benchmark and task authoring.
  • High attention to detail, creativity in finding what others missed, strong written communication, and ability to work independently on ambiguous, open-ended problems.
  • Capacity to engage reliably for approximately 35 hours per week.
Work Terms

This is a full-time W-2 employment position with Cincinnatus LLC, and the role places you within a leading AI lab as part of their extended workforce. The position is fully remote within the United States, with an expected commitment of approximately 35 hours per week. These roles are structured, role-based positions rather than freelance or project-only engagements, and involve close collaboration with client internal teams and integration into standard enterprise workflows.

Compensation

Hourly pay range: $60.00 to $90.00 per hour.

Eligibility
  • The position is remote within the United States and will be administered on a U.S. W-2 payroll, applicants must be eligible to work in the United States.
  • Cincinnatus LLC is an Equal Employment Opportunity employer, and reasonable accommodations are available for qualified individuals with disabilities throughout the application process.
Application process

Opportunities may be listed or discovered through partner channels, however Cincinnatus LLC administers hiring, onboarding, payroll, and benefits for this role. If you apply, expect standard employer-led onboarding and integration into the client team as a W-2 employee of Cincinnatus.

Related Jobs