About this role
Role Overview
This role operates at the frontier of reinforcement learning, building production-ready environments, training pipelines, data generation systems, and automated evaluation frameworks that translate research experiments into scalable systems. You will design complex, self-contained RL tasks, drive reproducible multi-component training processes, and help benchmark and validate model capabilities for deployment.
Key Responsibilities- Architect self-contained RL environments that model complex, real-world tasks, including reward functions, verifiers, and evaluation logic.
- Design and scale episode pipelines and multi-component training processes to enable reproducible experimentation.
- Build automated data generation systems, including synthetic data pipelines, to accelerate training cycles while maintaining data quality.
- Develop and integrate AI-driven evaluation and quality assurance systems for automated grading, validation, and feedback loops.
- Fine-tune and optimize open-source RL models using internally generated datasets and custom training strategies.
- Establish benchmarking frameworks to measure model capability, robustness, and data quality across tasks.
- Contribute to the release and analysis of evaluations on internal and external benchmark platforms, for example micro1 benchmarks or similar ecosystems.
- Deep experience in reinforcement learning, including environment design and training dynamics.
- Demonstrated track record building and scaling RL systems, experiment pipelines, or experimentation frameworks.
- Proficiency in automation and data generation, including synthetic data pipelines and scalable episode generation.
- Familiarity with automated evaluation systems, model validation, and quality assurance workflows.
- Experience fine-tuning and evaluating open-source ML models.
- Strong technical communication and writing skills, with the ability to document experiments and results clearly.
- Comfortable working in a fast-paced, research-driven, and highly collaborative environment.
- Preferred but not required, experience publishing benchmarks, evaluations, or research artifacts, and experience with large-scale RL experimentation infrastructure.
- Employment type: Full-time.
- Location: Remote, remote-first workforce.
- Role mixes research and production responsibilities and requires collaboration across engineering and research teams.
- Advertised total compensation range: $220, 000 to $500, 000 per year.
- National base salary range notice: $140, 000 to $180, 000 USD for this full-time position.
- All employees are eligible for equity compensation. Employees may also receive performance-based bonuses, subject to role and company policies.
- Benefits include up to 100% reimbursement for health insurance premiums, paid time off, a 401(k) plan with company match, and additional benefits for a remote-first workforce.
- All qualified applicants will receive consideration regardless of race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected by applicable law.
- Equity and bonus eligibility apply to employees as described in the Compensation section, and specific awards are subject to company policy.