About this role
Lead the strategy and hands-on execution for the infrastructure powering production AI systems. You will scale secure, observable, and resilient cloud platforms while shaping developer experience, reliability practices, and a high-performing infrastructure engineering organization.
Key Responsibilities- Own the architecture and evolution of multi-cloud infrastructure across AWS and GCP, balancing scalability, reliability, security, and cost.
- Lead, mentor, and grow an Infrastructure and Platform Engineering team with a culture of ownership, operational excellence, and continuous improvement.
- Build reproducible, version-controlled, auditable infrastructure using Terraform or an equivalent infrastructure-as-code approach.
- Design and improve CI/CD systems for fast, dependable, and secure software delivery.
- Establish observability across metrics, logs, traces, and alerting to identify and resolve production issues proactively.
- Drive reliability practices including incident response, on-call operations, SLOs, error budgets, disaster recovery, and blameless postmortems.
- Partner with Security and Engineering leadership to embed security by default and support compliance with ISO 27001, SOC 2, and CMMC.
- Reduce operational friction through automation and platform tooling that improves the developer experience.
- 8+ years building and operating production infrastructure, platform engineering, DevOps, or SRE systems.
- 3+ years leading engineering teams in high-growth environments.
- Deep AWS, GCP, or multi-cloud production experience.
- Strong experience with Terraform, Kubernetes, containers, and modern CI/CD platforms.
- Demonstrated ability to build highly available, observable, and resilient production systems.
- Strong knowledge of infrastructure security, compliance, and operational risk management.
- Experience scaling engineering organizations and infrastructure in fast-moving startup environments.
- Excellent communication skills and the ability to influence technical strategy with engineering and executive stakeholders.
- Experience supporting AI or ML infrastructure, large-scale data platforms, model training, inference, evaluation, or data-pipeline infrastructure.
- Experience operating regulated cloud environments, including FedRAMP, GovCloud, or CMMC Level 2.
- Contributions to platform engineering, open-source infrastructure, or developer productivity initiatives.
- Full-time, remote position.
- Listed compensation range: $350, 000 to $500, 000 per year.
- The stated national base-salary range is $230, 000 to $260, 000.
- Employees are eligible for equity compensation and may also receive performance-based bonuses.