Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

LLM Red Team Specialist for AI Model Evaluation

$60–$90/hr

Remote — US onlyUnited StatesContracttechnology
Apply Now

About this role

Role Overview

Find the subtle places where frontier language models appear competent but fail, and turn those discoveries into durable evaluation tasks. In this red-teaming role you will design and execute multi-step probes across coding, machine learning, and analysis domains to surface vulnerabilities, edge cases, and failure modes. Each probe is scoped to one to two days of focused work and feeds directly into benchmark development through rapid collaboration with researchers and task authors.

Key Responsibilities
  • Probe models: explore model behavior on coding, ML, and analytical challenges to identify where outputs are incorrect, inconsistent, or exploitable.
  • Design challenges: convert discovered weaknesses into well-crafted tasks that are difficult for models but fair and reproducible to grade.
  • Document findings: produce clear, evidence-backed writeups with reproducible steps, test cases, and evaluation notes.
  • Strengthen tasks: collaborate with task authors and researchers to close loopholes, remove shortcuts, and fix grading gaps.
  • Work as a team: share insights and iterate with other red-teamers and researchers to continuously improve the benchmark suite.
Qualifications
  • MSc or PhD in a STEM field, or equivalent practical experience in a research-focused domain involving data analysis and coding.
  • At least 1 year of experience in a research, research-engineering, security, or AI-evaluation role.
  • Proven ability to find vulnerabilities, edge cases, or failure modes in LLMs or ML systems via red teaming, adversarial testing, security research, or rigorous model evaluation.
  • Working proficiency in Python and Git, able to script probes and analyses independently.
  • Strong familiarity with LLM capabilities, limitations, and common evaluation techniques.
  • Preferred: prior experience in AI training, model evaluation, or benchmark and task authoring.
  • Mindset: high attention to detail, creativity in finding what others miss, strong written communication, and the ability to work independently on ambiguous, open-ended problems.
  • Availability to engage reliably for approximately 35 hours per week.
Work Terms

This is a full-time W-2 employment position administered by Cincinnatus LLC, with placement on a leading AI lab as part of their extended workforce. The role is fully remote within the United States and is structured as a role-based position rather than a project-based or freelance engagement. Typical work commitment is approximately 35 hours per week. Each individual probe or task generally requires one to two days of continuous focused effort and will span coding, experimentation, and analysis.

Employment, onboarding, payroll, benefits, and compliance are handled by Cincinnatus LLC. You may learn about this opportunity through partner listings or referral channels, but the employer of record and administrative handling will be Cincinnatus LLC.

Compensation

Hourly pay range: 60 to 90 per hour.

Eligibility

Position is fully remote within the United States. Cincinnatus LLC is an equal employment opportunity employer and provides reasonable accommodations for qualified individuals with disabilities. Roles placed through Cincinnatus are not freelance or gig-style engagements; they involve integration into client teams, regular collaboration, and adherence to enterprise workflows.

Related Jobs