Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

Software Engineer for AI Model Evaluation

$400,000–$800,000/yr

RemoteFull-timetechnology
Apply Now

About this role

Lead the design and evaluation of next-generation coding agents by creating benchmarks, measurement methodologies, datasets, and the tooling that enables rigorous, large-scale assessment and improvement of coding models.

Key Responsibilities
  • Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards.
  • Lead end-to-end research initiatives that measure and improve coding model performance across diverse software engineering tasks.
  • Develop high-quality datasets, golden examples, and evaluation protocols to enable reliable assessment of frontier coding systems.
  • Analyze model behavior and failure modes, identify systematic weaknesses, and translate findings into actionable improvements for training and evaluation.
  • Build tooling and infrastructure to support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
  • Establish and document best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
  • Collaborate with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
  • Contribute to technical reports, benchmark studies, and client-facing research deliverables that communicate model performance and insights.
Qualifications
  • Required skills: LLMs, coding, evaluation, AI evaluation, ML systems.
  • Strong software engineering background with expertise in Python, C++, or comparable programming languages.
  • Minimum 3 years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
  • Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
  • Strong analytical skills, with the ability to investigate model behavior and derive insights from complex technical systems.
  • Excellent written and verbal communication skills, including the ability to clearly articulate technical findings to diverse audiences.
  • Comfortable operating in fast-moving research environments with significant ambiguity and evolving priorities.
  • Preferred experience: working on frontier AI systems, coding agents, or model evaluation research; designing benchmarks or datasets for machine learning at scale; familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies.
  • Preferred evidence of impact: publications, open-source contributions, or demonstrated technical leadership.
Work Terms
  • Employment type: Full-time.
  • Location: Remote.
Compensation
  • Salary range: $400, 000 to $800, 000 per year.
Eligibility

This is a full-time remote position. The listing does not specify work authorization or visa sponsorship details, candidates should ensure they are able to work in a remote capacity under their own authorization.

Related Jobs