About this role
Lead hands-on research, evaluation, and data initiatives that improve the performance of frontier AI models and real-world AI systems. This role owns the path from experimental signal to defensible, evidence-based improvements in deployed systems.
Key Responsibilities- Own research and evaluation initiatives end to end, including problem framing, data design, quality calibration, and signal validation.
- Design ML-oriented data systems, including task definitions, annotation schemas, rubrics, incentives, and pipelines optimized for downstream model performance.
- Analyze model and system failures to identify root causes, edge cases, and opportunities for improvement.
- Convert ambiguous real-world behavior into structured evaluation frameworks and new data categories.
- Partner with researchers, domain experts, and operators to ensure experimental work produces clean, defensible research signal that translates into meaningful system improvements.
- Calibrate quality early with researchers and domain experts, then continuously raise the standard for signal quality.
- Iterate rapidly on evaluations, datasets, and feedback loops to improve system performance.
- Serve as a quality gate by challenging claims, pausing work, or changing scope when signal strength or data integrity is insufficient.
- Work with cross-functional and client-facing teams to communicate research progress through clear, credible, evidence-grounded narratives.
- Identify gaps in data or evaluation coverage and recommend where to invest, iterate, or stop based on learning and impact.
- Strong judgment about research signal quality and whether work is ready to be shared externally.
- Experience designing ML-oriented datasets, evaluation frameworks, and quality-assurance processes.
- Ability to translate messy real-world system behavior into structured research and evaluation opportunities.
- Comfort operating in ambiguity, with a strong sense of ownership and decisive action.
- Clear written and verbal communication skills, including explaining tradeoffs, limitations, and signal strength to technical and non-technical stakeholders.
- Demonstrated ability to work directly with experts through project kickoff, calibration, and iteration.
- A systems-level mindset focused on improving end-to-end model or agent performance.
- Experience with reinforcement learning environments, simulators, or feedback-driven training systems.
- Experience improving agentic AI systems or systems used in real-world workflows.
- Prior applied-research or production experience with direct impact on deployed systems.
- Experience designing evaluations for complex or real-world tasks.
- Familiarity with expert incentive design and engagement for high-stakes technical projects.
- Full-time, remote position.
- Stated annual compensation range: $600, 000 to $2, 000, 000.
- National base salary range: $180, 000 to $320, 000 per year.