About this role
Role Overview
Help ensure software engineering benchmark tasks are accurate, reproducible, and resistant to invalid solutions. You will evaluate repository level tasks used to train and assess advanced AI models, reviewing task quality, reference patches, test harnesses, grading integrity, and written feedback against defined rubrics.
Key Responsibilities
- Assess the quality, correctness, and reproducibility of software engineering benchmark tasks.
- Audit repository level tasks, reference patches, test runners, and Docker isolation.
- Identify answer leakage, reward hacking risks, and weaknesses in grading integrity.
- Provide clear, rubric based written feedback on task quality and evaluation readiness.
Qualifications
- At least 3 years of professional software engineering experience.
- Open source contribution or maintainer experience, such as merged pull requests or committer or maintainer roles.
- Strong ability to audit reference patches, test runners, and Docker isolation environments.
- Fluency in Python and at least one of Java, Go, TypeScript, or C++.
Preferred Qualifications
- Familiarity with SWE Bench Verified or similar repository based benchmarks.
- Maintainer experience with major Python open source projects, such as Django, Flask, scikit learn, SymPy, or pytest.
- Prior code review or task grading experience.
Work Terms
- Remote role open to candidates located in the United States.
- Hourly engagement.
Compensation
- $70 to $90 per hour.