Senior Platform Engineer for AI Model Training
$60–$130/hr
About this role
Apply deep cloud infrastructure and platform engineering expertise to design realistic, production-grade scenarios that train and evaluate next-generation AI systems. You will author Reinforcement Learning environments that test models on designing, deploying, troubleshooting, securing, scaling, and recovering distributed cloud systems. No prior AI domain experience is required, the role prioritizes hands-on, production ownership and engineering judgment.
Key Responsibilities- Create realistic cloud infrastructure tasks covering distributed systems, networking, security, scalability, and reliability.
- Build reproducible, containerized environments with valid golden reference solutions and intentionally defective variants for failure and resilience testing.
- Define measurable requirements across infrastructure configuration, deployed topology, and runtime behavior.
- Develop deterministic integration, load, security, failure-injection, deployment, and recovery tests.
- Debug environments, document technical decisions, and review tasks produced by other experts to ensure quality and correctness.
Required
- Senior-level experience in cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE, including personal ownership of a production platform.
- Core skills: Cloud Benchmark Task Authoring, Cloud and Distributed Systems Architecture, Production Infrastructure Ownership.
- Strong understanding of distributed systems, scalable APIs, queues, autoscaling, durable storage, and partial-failure scenarios.
- Practical experience with IAM, private networking, least-privilege access, and service-to-service security.
- Experience with observability, measurable SLOs, rolling deployments, rollback strategies, and disaster recovery.
- Ability to write infrastructure automation or testing tools and to debug containerized environments using a relevant programming language.
Preferred
- Experience with Terraform or OpenTofu.
- Experience on AWS, Azure, GCP, Kubernetes, or multi-cloud deployments.
- Experience building internal developer platforms, edge infrastructure, or shared platform services.
- Experience with chaos engineering, fault injection, local cloud emulators, or resilience testing.
- Experience creating technical evaluations, automated grading systems, or AI training/evaluation environments is helpful but not required.
- Role type, contractor. Location, remote.
- Work is task based, with experts responsible for completing reproducible tasks that meet project specifications.
- Experts must submit a minimum of tasks per week.
- Time to complete tasks will vary by the expert''s experience and workflow.
- Listed pay range: $60 to $130 per hour.
- Compensation model is output based, experts are paid per task that meets the project specifications.
- Minimum submission requirements apply, and payment depends on meeting the task acceptance criteria.
Assignments are provided on a task basis and depend on project needs and availability. The company vets and selects contributors through its talent identification process to scale expert contributions.
Application Process- Apply to the role and complete the screening questions.
- Complete an AI interview, approximately 30 minutes, which is reviewed by recruiters.
- Final review by the hiring manager.