Software Engineer for AI Code Evaluation and Benchmarking
from $50/hour
Remote — US onlyContracttechnology
Apply NowAbout this role
Role Overview
Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced engineers who enjoy code review, debugging, and applying software engineering judgment to improve model correctness and reliability.
Key Responsibilities- Review AI-generated code for correctness, efficiency, maintainability, and compliance with task requirements.
- Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
- Debug code, reproduce issues, and verify fixes across multiple programming environments.
- Evaluate model-generated explanations and reasoning for technical accuracy and soundness.
- Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
- Identify edge cases, failure modes, and areas where models struggle with software engineering problems.
- Document findings clearly and provide structured feedback to improve evaluation consistency and quality.
- Collaborate with project teams to establish and maintain quality standards and evaluation methodologies.
- Bachelor''s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
- Minimum 3 years of professional software engineering experience.
- Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
- Solid understanding of data structures, algorithms, software design principles, and debugging methodologies.
- Experience performing code reviews and evaluating code quality in production or large-scale codebases.
- Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
- Familiarity with version control systems such as Git and with modern software development workflows.
- Strong written communication skills and attention to detail.
- Experience with AI or ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.
- Experience evaluating AI-generated code, creating benchmarks, or assessing software quality is highly preferred.
- Engagement type: Contractor assignment, no medical or paid leave provided.
- Minimum commitment: at least 4 hours per day and at least 20 hours per week, with a required daily overlap of 4 hours aligned to PST.
- Contract length: 1 month, expected start date is next week.
- Work location: Remote, United States only.
- Perks: fully remote work and the opportunity to contribute to cutting-edge AI coding evaluation projects.
- Applicants must be located in the United States. This role is open to US-based candidates only.
Candidates will complete an online automated coding assessment covering Python and a Docker-based test identified as RHLF as part of the selection process.