Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

Data Scientist for AI Model Evaluation

$60–$90/hr

Remote — US onlyUnited StatesContracttechnology
Apply Now

About this role

Role Overview

Help build next-generation agentic evaluation benchmarks for frontier AI models by acting as a ground-truth expert in data science and quantitative analysis. You will design and execute realistic, research-style analysis tasks that test model capabilities, verify statistical claims, and produce clear, reproducible notebooks that drive researcher decisions. Tasks typically require one to two days of continuous, focused effort and span data cleaning, statistical analysis, interpretation, and written reporting. You will operate in a close feedback loop with researchers to pinpoint where advanced models fall short on rigorous analytical work.

Key Responsibilities
  • Design complex, realistic data-analysis tasks that simulate actual research work, including cleaning messy datasets, defining fair comparisons between methods, and specifying success criteria.
  • Author reproducible analyses and reports in Jupyter Notebooks or Google Colab that a researcher can follow and act on.
  • Construct comparisons between analytical approaches, for example comparing two anomaly-detection algorithms on a dataset, calculating correlations, and recommending a preferred method backed by spot checks.
  • Perform manual spot checks and statistical validation, interpret results carefully, and summarize findings clearly enough to inform research decisions.
  • Evaluate how models perform on your tasks, confirming whether reported statistics and conclusions hold up under scrutiny.
  • Coordinate with researchers and other subject-matter experts to align evaluation standards and maintain consistency across tasks.
Qualifications
  • MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain.
  • At least 1 year of experience in a research, research-engineering, or heavy data-analysis role.
  • Deep, hands-on skills in data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results.
  • Proficiency using Jupyter Notebooks or Google Colab for analysis and reporting.
  • Working proficiency in Python, including libraries such as pandas and NumPy, and familiarity with Git.
  • Strong written communication skills, capable of explaining analytical findings to decision-makers.
  • Experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • A perfectionist mindset, high attention to detail, creativity in task design, and the ability to work independently on ambiguous, open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.
Work Terms
  • Employment type: full-time, W-2 employment with Cincinnatus LLC, serving as employer of record for placement at a leading AI lab as part of their extended workforce.
  • Location: fully remote within the United States.
  • Typical weekly commitment: approximately 35 hours per week.
  • Task cadence: individual evaluation tasks generally represent one to two days of continuous, focused effort.
  • Role structure: these are structured, role-based positions rather than freelance or one-off project engagements, and involve close collaboration with client teams and integration into enterprise workflows.
Compensation
  • Hourly rate: $60.00 to $90.00 per hour, W-2.
  • Payroll, benefits, and related employment administration are provided by Cincinnatus LLC.
Eligibility
  • Candidates must be able to accept W-2 employment and work remotely within the United States.
  • Reliable availability to commit approximately 35 hours per week is required.
Application Process

Apply via the job posting channel where you found this opportunity. Employment, onboarding, payroll, and benefits for selected candidates will be administered by Cincinnatus LLC, and selected hires will be placed to work with the client research team as part of their extended workforce.

Related Jobs