About this role
Role Overview
Help shape how production LLM agents are built, evaluated, operated, and adopted inside organizations. This opportunity focuses on engineers with firsthand experience delivering agents that users rely on and understanding the practical tradeoffs behind reliable, useful systems.
Key Responsibilities
- Bring experience with production agents, including systems that run as scheduled jobs, perform long-running work, or operate in cloud sandboxes.
- Apply approaches for determining whether an agent improved or regressed after changes.
- Share insight into internal assistants, including adoption patterns and why some employees may not use them.
- Discuss agent infrastructure such as company-data integrations, shared organizational memory, reusable skills and playbooks, tool and MCP interfaces, and observability for agent activity and cost.
Qualifications
- Have shipped an agent that real users depended on and taken responsibility when it failed.
- Bring concrete experience and detailed examples involving agent reliability, evaluation, adoption, and real-world engineering tradeoffs.
Work Terms
- Remote, per-task engagement.
- The application process begins with a short conversational AI interview, with no coding exercise or take-home assignment. It is intended to assess how you think through reliability, evaluation, adoption, and system tradeoffs.
- Strong candidates will be invited to a live 30-minute conversation with the team.
- Every completed submission is reviewed.
Compensation
- Live conversation compensation is $100 to $500 per completed call, based on depth of experience.