About this role
Role Overview
This position focuses on enhancing the safety of AI systems through rigorous testing and evaluation. As part of a dedicated red team, you will probe AI models with adversarial inputs to identify vulnerabilities and generate data that improves AI safety for users.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structured methodologies by following taxonomies, benchmarks, and playbooks to ensure consistent testing.
- Document findings reproducibly, producing reports, datasets, and attack cases that customers can act upon.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- A curious and adversarial mindset, with a natural inclination to push systems to their limits.
- Structured approach to testing, utilizing frameworks or benchmarks rather than random methods.
- Strong communication skills to clearly explain risks to both technical and non-technical stakeholders.
- Adaptability to thrive in a dynamic environment with varying projects and customers.
- Experience with adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, and model extraction.
- Background in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
- Knowledge of socio-technical risks, including harassment/disinformation probing and conversational AI testing.
- Creative probing skills in psychology, acting, or writing that foster unconventional adversarial thinking.
- Identifying vulnerabilities that automated tests overlook.
- Delivering reproducible artifacts that enhance the robustness of customer AI systems.
- Expanding evaluation coverage by testing more scenarios and reducing surprises in production.
- Building trust with customers in the safety of their AI systems through thorough adversarial probing.
Gain valuable experience in human data-driven AI red teaming while contributing to the development of more robust, safe, and trustworthy AI systems.