About this role
Role Overview
Lead the design, delivery, and day to day operations of the infrastructure that powers production AI systems. This hands-on leadership role owns multi cloud architecture, reliability, security, and the developer platform, and combines long term strategy with incident response and engineering mentorship to scale a high performing infrastructure organization.
Key Responsibilities- Own the architecture and evolution of multi cloud infrastructure across AWS and GCP, optimizing for scalability, reliability, security, and cost.
- Lead, hire for, and grow a high performing Infrastructure and Platform Engineering team, establishing an engineering culture of ownership, operational excellence, and continuous improvement.
- Author and maintain infrastructure as code, primarily with Terraform or equivalent tools, to ensure reproducible, version controlled, auditable environments.
- Design, implement, and improve CI/CD systems to enable fast, reliable, and secure software delivery.
- Build and operate comprehensive observability including metrics, logs, traces, and alerting to proactively detect and resolve production issues.
- Define and drive reliability practices, including incident response, on call operations, service level objectives and error budgets, disaster recovery, and blameless postmortems.
- Partner closely with Security and Engineering leadership to embed security by default and support compliance with frameworks such as ISO 27001, SOC 2, and CMMC.
- Reduce operational friction and improve developer experience through automation and platform tooling.
- Participate directly in solving production incidents and mentor engineers across the stack.
- Minimum 8 years building and operating production infrastructure, platform engineering, DevOps, or SRE systems.
- At least 3 years leading engineering teams in high growth environments.
- Deep, demonstrable experience with AWS and GCP in production, including multi cloud architecture patterns.
- Strong experience with Terraform, Kubernetes, containers, and modern CI/CD platforms.
- Proven track record building highly available, observable, and resilient production systems.
- Solid understanding of infrastructure security, compliance, and operational risk management, with experience partnering with security teams.
- Excellent communication skills, able to influence technical strategy across engineering and executive stakeholders.
Preferred Experience
- Experience supporting AI and ML infrastructure or large scale data platforms, including model training, inference, evaluation, or data pipeline infrastructure.
- Experience operating regulated cloud environments such as FedRAMP, GovCloud, or CMMC Level 2.
- Contributions to platform engineering, open source infrastructure, or developer productivity initiatives.
- Employment type: Full time.
- Location: Remote.
- Team: Core Infrastructure/Platform Engineering.
- Primary compensation range for this role is 350000 to 500000 yearly.
- Additional pay notice: the national base salary range for this full time position is 230000 to 260000.
- All employees are eligible for equity compensation, and employees may also receive performance based bonuses.
- This is a full time employee role. All employees are eligible for equity and may be eligible for performance based bonuses as noted above.