Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

Member of Technical Staff, Coding Research

$400,000–$800,000/yr

RemoteFull-timetechnology
Apply Now

About this role

Role Overview

Advance the evaluation and development of frontier coding agents by designing the benchmarks, methodologies, and data systems used to measure and improve next-generation coding models. This role combines AI research, software engineering, and model evaluation.

Key Responsibilities

  • Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methods, rubrics, and quality standards.
  • Lead end-to-end research initiatives to measure and improve coding-model performance across a range of software engineering tasks.
  • Develop high-quality datasets, golden examples, and evaluation protocols for reliable assessment of frontier coding systems.
  • Analyze model behavior and failure modes, identify systematic weaknesses, and translate findings into improvements for training and evaluation.
  • Build tooling and infrastructure for large-scale experimentation, data generation, review workflows, and evaluation pipelines.
  • Establish rigorous, reproducible best practices for coding-agent assessment and measurement quality.
  • Partner with researchers, engineers, and applied AI teams to design experiments and assess emerging model capabilities.
  • Contribute to technical reports, benchmark studies, and client-facing research initiatives that communicate model performance and insights.

Qualifications

  • Strong software engineering background and expertise in Python, C++, or comparable programming languages.
  • At least 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical discipline.
  • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
  • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
  • Ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
  • Strong analytical skills for investigating model behavior and deriving insights from complex technical systems.
  • Excellent written and verbal communication skills, including the ability to explain technical findings to varied audiences.
  • Comfort working in a fast-moving research environment with ambiguity and evolving priorities.

Preferred Qualifications

  • Experience with frontier AI systems, coding agents, or model-evaluation research.
  • Interest in how data, evaluations, and feedback mechanisms influence model capabilities.
  • A record of independently driving ambiguous technical or research projects from concept through execution.
  • Experience designing benchmarks or datasets for machine learning systems at scale.
  • Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methods.
  • Publications, open-source contributions, or demonstrated technical leadership.

Work Terms

  • Full-time, remote role.

Compensation

  • Annual compensation range: $400, 000 to $800, 000.

Related Jobs

More like this