Key details
- Role type
- Contract
- Compensation
- $70–$90/hr
- Work arrangement
- Remote
- Category
- technology
- Confirmed requirements
- 4
About this role
Role Overview
Evaluate Neuron Kernel Interface, NKI, development tasks used to train and assess advanced AI models. You will review CUDA-to-NKI migrations, Trainium-focused performance optimizations, and cross-platform numerical correctness, then deliver clear written feedback using defined rubrics.
Key Responsibilities
- Assess NKI kernel-development tasks for quality, correctness, and suitability for AWS Trainium and Inferentia2 hardware.
- Evaluate the fidelity of CUDA-to-NKI migrations.
- Review Trainium-specific performance optimization quality.
- Evaluate numerical-correctness standards across GPU and Trainium platforms.
- Provide clear, rubric-based written feedback.
Qualifications
- At least 2 years of hands-on experience developing or optimizing NKI kernels for AWS Trainium or Inferentia2 hardware.
- Strong knowledge of tile-based computation, SBUF, PSUM, and HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration.
- Experience evaluating CUDA-to-NKI migration quality.
- Familiarity with Trainium performance profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks.
- Experience establishing or assessing cross-platform numerical-correctness standards, including GPU versus Trainium accumulation order, rounding behavior, and mixed-precision semantics.
Preferred Qualifications
- Experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel-library contributions.
- Prior CUDA or Triton kernel-development experience.
- Familiarity with NeuronCore-v2 architecture, on-chip SRAM topology, and FP32, BF16, FP8, and INT8 data types.
- Experience benchmarking machine-learning training workloads on Trn1 or Trn2 instances.
Work Terms
- Remote role, open to candidates located in the United States.
- Hourly engagement.
Compensation
$70 to $90 per hour.
What to prepare before applying
- 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS Trainium/Inferentia2 hardware
- Strong understanding of NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration
- Demonstrated experience assessing CUDA→NKI migration quality
- Familiarity with Trainium-specific performance profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.