Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

GPU Kernel Developer for AI Model Evaluation

$70–$90/hr

RemoteRemote — United StatesContracttechnology
Apply Now

Key details

Role type
Contract
Compensation
$70–$90/hr
Work arrangement
Remote
Category
technology
Confirmed requirements
4

About this role

Role Overview

Assess GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. You will review task quality, numerical correctness, completeness, benchmarking fairness, scope, and compilation and runtime validity across a range of kernel-development scenarios, then deliver clear written feedback using defined evaluation rubrics.

Key Responsibilities

  • Evaluate GPU and accelerator kernel tasks for quality, correctness, completeness, and appropriate scope.
  • Review numerical-validation approaches, performance-benchmarking fairness, and compilation and runtime validity.
  • Provide precise, rubric-based written feedback across diverse kernel task types.

Qualifications

  • 3+ years of hands-on experience developing, optimizing, or verifying GPU or accelerator kernels in at least two of CUDA, Triton, NKI, or Pallas for JAX.
  • Strong knowledge of kernel numerical-correctness standards, including absolute, relative, and ULP tolerances and selection of reference implementations.
  • Experience profiling and benchmarking performance with tools such as Nsight, NCU, roofline analysis, or framework-native profilers.
  • Familiarity with compilation and runtime failure modes, including driver mismatches, out-of-memory errors, launch-configuration errors, shape or stride mismatches, and autotuning failures.
  • Experience with at least three kernel task types: specification-based generation, cross-framework translation or lowering, hardware-target migration, debugging, performance optimization, or operator fusion.

Preferred Qualifications

  • Experience in both NVIDIA GPU environments, including CUDA or Triton, and custom accelerator environments, including NKI, Pallas, or TPU.
  • Background in compiler engineering, MLIR, or intermediate-representation lowering.
  • Knowledge of memory-hierarchy optimization, including shared-memory tiling, register pressure, bank conflicts, and coalescing patterns.
  • Contributions to kernel libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls.

Work Terms

  • Remote role open to candidates located in the United States.
  • Hourly engagement.

Compensation

  • $70 to $90 per hour.

What to prepare before applying

  1. 3+ years of hands-on experience developing, optimizing, or verifying GPU/accelerator kernels in at least two of CUDA, Triton, NKI, or Pallas (JAX).
  2. Strong understanding of numerical-correctness criteria for kernels, including absolute, relative, and ULP tolerances and reference-implementation selection.
  3. Demonstrated experience with performance profiling and benchmarking, such as nsight, ncu, roofline analysis, or framework-native profilers.
  4. Familiarity with common compilation and runtime failure modes, including driver mismatches, OOM, launch-configuration errors, shape/stride mismatches, and autotuning failures.

These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.

Related Jobs

More like this