About this role
Role Overview
Assess Kubernetes tasks for quality, technical correctness, and production readiness as they are used to train and evaluate advanced AI models. You will review cluster-operations scenarios, manifest accuracy, and failure-mode troubleshooting, then deliver clear written feedback using defined rubrics.
Key Responsibilities
- Evaluate Kubernetes tasks for quality, correctness, and production readiness.
- Review cluster-operations scenarios, Kubernetes manifests, and troubleshooting approaches.
- Provide clear, rubric-based written feedback on task quality and technical accuracy.
Qualifications
- At least 3 years of hands-on production Kubernetes experience with EKS, GKE, AKS, or self-managed environments.
- Strong knowledge of cluster internals, including CNI, DNS, ingress, PV/PVC storage, RBAC, and operational failure modes such as CrashLoopBackOff, OOMKilled, scheduling, and eviction.
- Experience authoring and reviewing manifests or Helm charts, and debugging live cluster incidents.
- Proficiency in Go, Python, or TypeScript.
Preferred Qualifications
- CKA or CKAD certification.
- Experience with service meshes, autoscaling, and observability tools, including Istio, HPA, Prometheus, and Grafana.
- Prior experience in SRE, platform engineering, or task grading.
Work Terms
- Remote role open to candidates located in the United States.
- Hourly engagement.
Compensation
$70 to $90 per hour.