About this role
Role Overview
Lead hands-on technical work at the intersection of applied AI research, data design, and real-world model systems. You will own the path from rigorous evaluation and failure analysis to iterative improvements in deployed models and agents, partnering closely with researchers, domain experts, and operators to produce defensible research signal.
Key Responsibilities
- Own research and evaluation initiatives end to end, including problem framing, data design, quality calibration, and signal validation.
- Design machine-learning-oriented data systems, including task definitions, annotation schemas, rubrics, incentives, and pipelines that support downstream model performance.
- Analyze model and system failures to identify root causes, edge cases, and improvement opportunities.
- Convert ambiguous real-world behavior into structured evaluation frameworks and new data categories.
- Collaborate with researchers and domain experts to calibrate quality early and continuously raise the standard for useful signal.
- Rapidly iterate on evaluations, datasets, and feedback loops to improve system performance.
- Serve as a quality gate by blocking claims, pausing work, or requiring scope changes when data integrity or signal strength is insufficient.
- Partner with cross-functional and client-facing teams to communicate research progress through clear, evidence-based narratives.
- Identify gaps in data and evaluation coverage, then recommend where to invest, iterate, or stop based on impact and learnings.
Qualifications
- Strong judgment about research-signal quality and whether work is ready to be shared externally.
- Experience designing ML-oriented datasets, evaluation frameworks, and quality-assurance processes.
- Ability to turn messy real-world system behavior into structured research and evaluation opportunities.
- Comfort operating amid ambiguity, with decisive ownership and action.
- Clear written and verbal communication when explaining tradeoffs, limitations, and signal strength to technical and non-technical stakeholders.
- Demonstrated ability to work directly with experts during project kickoff, calibration, and iteration.
- A systems-level mindset focused on end-to-end model or agent performance rather than isolated components.
Preferred Experience
- Reinforcement learning environments, simulators, or feedback-driven training systems.
- Improving agentic systems or AI systems used in real-world workflows.
- Applied research or production work with direct impact on deployed systems.
- Evaluation design for complex or real-world tasks.
- Expert incentive design and engagement on high-stakes technical projects.
Work Terms
- Full-time remote position.
Compensation
- Stated compensation range: $600, 000 to $2, 000, 000 per year.
- National base salary range for this full-time role: $180, 000 to $320, 000 per year.