Software Engineer for AI Task Curation
$60–$90/hr
About this role
Design and deliver realistic, multi-step software engineering challenges used to evaluate frontier generative AI coding agents. You will create tasks that represent one to two days of focused engineering work across Python implementation, environment and tooling setup, debugging, and documentation, and you will collaborate closely with researchers to identify precisely where and why advanced models fail.
Key Responsibilities- Design tasks that simulate real research-engineering problems, chosen to push current AI coding agents beyond their reliable capabilities.
- Implement clear reference solutions in Python, including necessary setup, checks, and verifiable outputs.
- Use AI coding assistants as part of your workflow, observing how they perform and where they break down.
- Review and refine tasks authored by other experts, providing feedback on clarity, correctness, and difficulty.
- Analyze agent attempts on your tasks and report concrete failure modes to the research team so they can improve evaluations and model understanding.
- MSc or PhD in computer science or another STEM field, or equivalent practical experience in a research-heavy role that required substantial coding and data analysis.
- At least 1 year of experience in research, research engineering, or software engineering.
- Strong hands-on Python scripting and debugging ability, with attention to clean, readable code.
- Familiarity with version control using Git, common IDEs, and standard software development workflows.
- Experience with AI coding assistants, prompt engineering, or agent workflows is preferred.
- Prior experience in AI training, model evaluation, or benchmark and task authoring is preferred.
- High attention to detail, creative problem design skills, strong written communication, and the ability to work independently on ambiguous, open-ended problems.
- Availability to engage reliably for approximately 35 hours per week.
- Employment type: hourly, W-2 employment with Cincinnatus LLC.
- Placement: role may be placed at a leading AI lab as part of an extended workforce integrated into a client team.
- Schedule: approximately 35 hours per week, fully remote within the United States.
- Engagement model: role-based employment, not a project-by-project freelance engagement; includes integration into standard enterprise workflows and close collaboration with client teams.
- Onboarding, payroll, benefits, and employment administration are handled by Cincinnatus LLC.
- Pay rate: $60.00 to $90.00 per hour.
- This position is fully remote within the United States and is employed on a W-2 basis by Cincinnatus LLC, which administers payroll and benefits.
- Opportunities may be advertised through recruiting platforms, but employment, onboarding, payroll, and benefits are administered by Cincinnatus LLC.
Apply using the job posting or recruiting channel where you discovered this opportunity. Applications will be processed by the hiring team, and final employment, onboarding, payroll, and benefits will be administered by Cincinnatus LLC.