About this role
Role Overview
Help advance frontier large language models by creating and evaluating high-quality ML systems training data. This hands-on infrastructure role focuses on GPU kernels, performance analysis, accelerated and distributed workload debugging, and high-throughput LLM serving within a leading AI research environment.
Key Responsibilities
- Create challenging, domain-relevant MLOps and ML systems tasks across GPU kernels, performance profiling, debugging, and inference serving, then produce accurate, well-structured solutions.
- Evaluate tasks and technical solutions, delivering clear written feedback that meets rigorous review standards.
- Help research and engineering teams address knowledge gaps and improve model performance in ML systems, training infrastructure, and framework-level topics.
- Develop detailed guidelines, rubrics, and evaluation frameworks for kernel optimization, profiler-output interpretation, distributed-systems reasoning, and serving throughput and latency trade-offs.
- Partner with subject matter experts to maintain consistent, accurate training data.
Qualifications
- 2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering.
- Experience in at least one of the following areas, with experience across multiple areas strongly preferred: custom GPU kernel development or optimization using CUDA, Triton, or Pallas; performance profiling and trace analysis using Kineto, torch.profiler, Nsight, or XLA or JAX profiler; debugging distributed or accelerator-bound workloads; or large-scale LLM serving using vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching.
- Production experience with JAX and/or PyTorch. Depth in custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler work, or graph-level work is a strong plus.
- Familiarity with A100, H100, B200, or TPU accelerators and the ability to assess throughput, latency, and memory trade-offs.
- Demonstrable career progression, strong written communication, and the ability to explain complex technical decisions clearly.
Work Terms
- United States-based, full-time W-2 hourly employment.
- 40 hours per week during weekdays are required. Candidates must have no conflicting engagements.
- Work may include placement with a leading AI lab as part of its extended workforce.
Compensation
- $90 to $120 per hour.
Equal Opportunity
Employment decisions are made without discrimination based on race, religion, color, national origin, sex, pregnancy, childbirth, reproductive health decisions, related medical conditions, sexual orientation, gender identity or expression, age, protected veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic.