About this role
Role Overview
Create original, executable scientific-computing evaluation problems that challenge leading AI models. This remote, part-time engagement focuses on translating advanced mathematics research and real-world computational scenarios into rigorous coding tasks and objective grading standards.
Key Responsibilities
- Develop original, executable research problems for scientific AI evaluation.
- Use a published paper, Kaggle dataset, open-source repository, or a self-designed scenario as source material.
- Write scientific coding prompts based on the selected material.
- Create grading criteria that clearly define correct answers.
- Test and calibrate tasks against frontier AI models; tasks are finalized when strong models fail more often than they succeed.
Qualifications
- PhD in mathematics, applied mathematics, computational mathematics, or a closely related field.
- Demonstrated, coding-focused depth in at least two of these areas: numerical linear algebra, computational mechanics, and computational finance.
- Working proficiency in Python or R for scientific computing.
- Comfort using Git and GitHub, and running code in Docker. Work is completed through a pull-request workflow with automated quality checks.
Preferred Qualifications
- Publications in peer-reviewed journals.
- Prior scientific software development or research engineering experience.
Work Terms
- Remote hourly engagement.
- Six-week project, with an immediate start date.
- Part-time commitment of at least 20 hours per week.
Compensation
- $70 per hour.
Application Process
- Submit your resume and completed application form.
- Complete a 25-minute conversational interview covering your background, experience, and motivations.
- Receive follow-up within a few days regarding next steps and onboarding.