Computational STEM Researcher for AI Model Evaluation
$60–$90/hr
About this role
Role Overview
Help build next-generation evaluation benchmarks for frontier AI models by turning real research practice into complex, multi-step challenges. Your expertise in experimental design, hypothesis testing, coding, data analysis, and rigorous conclusions will shape tasks that require one to two days of focused work and remain difficult for current models to complete reliably.
Key Responsibilities
- Design engaging, multi-step research tasks grounded in study design, hypothesis testing, and evaluation of results.
- Complete and document solutions to your own tasks using Python and notebook environments, with the rigor expected in professional research.
- Define clear standards that distinguish sound scientific reasoning from plausible but flawed reasoning.
- Review model responses and identify methodological, analytical, and reasoning errors that an experienced researcher would catch.
- Collaborate closely with researchers and subject-matter experts to keep evaluations accurate and consistent.
Qualifications
- MSc or PhD in a STEM discipline, a computational social science or humanities discipline, or equivalent practical experience in a research-intensive domain involving data analysis and coding.
- At least 1 year of experience in an active research role in academia, industry, or a national laboratory.
- Meaningful computational research experience, including Python-based analysis, simulation, modeling, or data pipelines.
- Strong knowledge of experimental design, hypothesis testing, and rigorous evaluation of results.
- Working familiarity with Git, IDEs, and Jupyter or Colab notebook environments.
- Experience with AI training, model evaluation, or benchmark or task authoring is preferred.
- Exceptional attention to detail, creativity in task design, strong written communication, and the ability to independently resolve ambiguous, open-ended problems.
Work Terms
- Full-time W-2 employment, with placement on an extended workforce team supporting a leading AI lab.
- Fully remote role, available only to candidates located in the United States.
- Approximately 35 hours per week, with reliable availability required.
- Employment, payroll, benefits, and compliance are administered by the employer of record.
Compensation
$60 to $90 per hour.
Application Process
Candidates may discover this opportunity through an online platform. If selected, employment onboarding, payroll, benefits, and related administration are handled by the employer of record.
Equal Employment Opportunity
Employment decisions are made without discrimination based on race, religion, color, national origin, sex, including pregnancy, childbirth, reproductive health decisions, related medical conditions, sexual orientation, gender identity, gender expression, age, protected veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans throughout the application process.