About this role
Apply advanced mathematical reasoning, computational thinking, and clear written explanations to improve and evaluate large language models and other AI systems. You will design challenging math problems, produce rigorous solutions and computational implementations, and review model outputs to identify errors and gaps. Work contributes to both frontier AI research and to projects that help enterprises deploy reliable AI.
Key Responsibilities- Contribute to research and product work that accelerates frontier AI, including supplying high quality math content, designing evaluation pipelines, and collaborating with expert researchers.
- Design original, challenging mathematics problems that probe multi step reasoning, abstract concepts, and proof based thinking for large language models.
- Solve problems independently and produce detailed, logically organized solutions with clear justifications suitable for human and model evaluation.
- Review and annotate model generated solutions, identify mathematical mistakes or missing arguments, and provide precise corrections and feedback.
- Help define new evaluation benchmarks derived from mathematics curricula spanning early undergraduate to PhD level topics.
- Design precise, closed ended prompts for evaluation tasks and write reliable Python code to implement computational checks, using approved scientific libraries, and verify numerical answers.
- Work on theorem prover tasks using the Lean proof assistant, translate problems and proofs into formal language, and ensure formal proofs compile and verify correctly.
- Strong foundation in mathematics at a level comparable to engineering entrance exams and graduate or PhD programs, with the ability to tackle multi step and proof oriented problems.
- Proven ability to break down complex mathematical concepts into simple, clear explanations, using plain language, visuals, and examples when appropriate.
- Experience writing reliable Python code for numerical or symbolic computation, and familiarity with common scientific libraries is preferred.
- Experience or interest in formal verification with Lean is a plus.
- Excellent structured written communication, attention to logical detail, and the ability to provide constructive, annotation style feedback.
- Creative and lateral thinking, good research and analytical skills, and the ability to work independently in a remote environment.
- Technical requirements: a desktop or laptop with a dependable internet connection.
- Engagement type: Contractor assignment, freelancer status. This role does not include medical or paid leave.
- Time commitment options: 20 hours per week, 30 hours per week, or 40 hours per week.
- Minimum commitment: at least 4 hours per day and a minimum of 20 hours per week.
- Scheduling requirement: maintain 4 hours of overlap with Pacific Time for collaboration.
- Work location: fully remote.
- Contract extensions: possible based on performance and project needs.
Compensation details are not specified in the job description.
Eligibility- Candidates pursuing a Master’s, PhD, or postdoctoral work in Mathematics, Applied Mathematics, Statistics, or a closely related field are eligible and encouraged to apply.
Eligible candidates are encouraged to submit an application describing relevant math background, examples of problem solving or teaching materials if available, and any experience with Python or Lean. Applications will be reviewed for fit with the contractor assignments and project needs.