About this role
Evaluate and strengthen advanced AI coding agents by applying hands-on machine learning engineering expertise to realistic technical workflows. This remote opportunity centers on assessing model-generated solutions across training, inference, MLOps, and LLM application work.
Key Responsibilities- Use advanced AI coding agents to complete and evaluate complex machine learning and AI engineering tasks.
- Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications.
- Identify bugs, edge cases, performance issues, and failure modes.
- Compare outputs from multiple advanced models and assess their respective strengths and weaknesses.
- Apply professional engineering judgment through structured technical assessments that help improve AI coding models.
- At least 2 years of professional machine learning engineering experience.
- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
- Regular experience using AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
- Ability to evaluate model-generated machine learning implementations and technical tradeoffs.
- Experience deploying ML systems to production is preferred.
- Remote, hourly engagement.
- Sprint-based work is scheduled in 12- to 24-hour stretches based on client requirements.
- $85 per hour.
- $400 for each accepted task.
- Typical tasks take approximately 2 to 3 hours after ramp-up.
- Payment is tied to accepted work.