Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

Software Engineer, AI Benchmark Task Curator

$60–$90/hr

Remote — US onlyUnited StatesContracttechnology
Apply Now

About this role

Design demanding, real-world software engineering challenges that help evaluate the limits of frontier AI coding agents. You will create tasks that reflect the complex work research engineers encounter, then collaborate with researchers to understand where AI agents succeed and fail.

Role Overview

Each task is designed to require one to two days of focused work and may combine Python implementation, environment and tooling setup, debugging, and clear documentation. This is a research-focused engineering opportunity supporting the development of advanced AI models.

Key Responsibilities
  • Create realistic, multi-step software engineering tasks that challenge current AI coding agents.
  • Implement reference solutions in Python, including the setup and validation checks needed to produce a clear, verifiable result.
  • Use AI coding assistants in daily work and assess where they are effective or fall short.
  • Review tasks created by other experts and provide feedback on clarity, correctness, and level of difficulty.
  • Analyze AI-agent attempts and help researchers identify the causes of task failures.
Qualifications
  • Master''s or PhD in computer science or another STEM discipline, or equivalent practical experience in a research-intensive field requiring substantial coding and data analysis.
  • At least 1 year of experience in research, research engineering, or software engineering.
  • Strong hands-on Python scripting and debugging ability, with attention to clean, readable code.
  • Comfort using Git, IDEs, and standard software development workflows.
  • Experience with AI coding assistants, prompt engineering, or agent workflows is preferred.
  • Experience with AI training, model evaluation, or benchmark or task authoring is preferred.
  • Exceptional attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
Work Terms
  • Fully remote role, available to candidates located in the United States.
  • W-2 employment in a structured role-based position, with potential placement on an AI lab''s extended workforce.
  • Approximately 35 hours per week, with reliable availability required.
  • This is not a freelance or project-based engagement. The role involves collaboration with internal client teams and standard enterprise workflows.
Compensation

$60 to $90 per hour.

Application and Employment

This opportunity may be found through a job platform. Employment, onboarding, payroll, benefits, and compliance are administered by the employer of record.

Equal Opportunity

Qualified applicants are considered without discrimination based on race, religion, color, national origin, sex, pregnancy or related medical conditions, sexual orientation, gender identity or expression, age, veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans during the application process.

Related Jobs