Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

Software Engineer for AI Code Evaluation and Benchmarking

from $50/hour

Remote — US onlyContracttechnology
Apply Now

About this role

Role Overview

Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced engineers who enjoy code review, debugging, and applying software engineering judgment to improve model correctness and reliability.

Key Responsibilities
  • Review AI-generated code for correctness, efficiency, maintainability, and compliance with task requirements.
  • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
  • Debug code, reproduce issues, and verify fixes across multiple programming environments.
  • Evaluate model-generated explanations and reasoning for technical accuracy and soundness.
  • Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
  • Identify edge cases, failure modes, and areas where models struggle with software engineering problems.
  • Document findings clearly and provide structured feedback to improve evaluation consistency and quality.
  • Collaborate with project teams to establish and maintain quality standards and evaluation methodologies.
Qualifications
  • Bachelor''s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
  • Minimum 3 years of professional software engineering experience.
  • Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
  • Solid understanding of data structures, algorithms, software design principles, and debugging methodologies.
  • Experience performing code reviews and evaluating code quality in production or large-scale codebases.
  • Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
  • Familiarity with version control systems such as Git and with modern software development workflows.
  • Strong written communication skills and attention to detail.
  • Experience with AI or ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.
  • Experience evaluating AI-generated code, creating benchmarks, or assessing software quality is highly preferred.
Work Terms
  • Engagement type: Contractor assignment, no medical or paid leave provided.
  • Minimum commitment: at least 4 hours per day and at least 20 hours per week, with a required daily overlap of 4 hours aligned to PST.
  • Contract length: 1 month, expected start date is next week.
  • Work location: Remote, United States only.
  • Perks: fully remote work and the opportunity to contribute to cutting-edge AI coding evaluation projects.
Eligibility
  • Applicants must be located in the United States. This role is open to US-based candidates only.
Evaluation Process

Candidates will complete an online automated coding assessment covering Python and a Docker-based test identified as RHLF as part of the selection process.

Related Jobs