About this role
Role Overview
Use adversarial testing to probe conversational AI models and agents, surface vulnerabilities, and produce reproducible attack artifacts that help customers harden their systems. This role focuses on text-based red teaming across sensitive topics such as bias, misinformation, and harmful behavior, with options to work on higher-sensitivity projects under clear guidelines and wellness support.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Review AI outputs that touch on sensitive topics, identify failures, and document harmful or risky behaviors
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Follow taxonomies, benchmarks, and playbooks to keep testing systematic and consistent
- Document reproducibly, delivering reports, datasets, and attack cases customers can act on
- Opt in to higher-sensitivity projects when appropriate, with topics communicated in advance and wellness resources available
- Prior red teaming experience, in AI adversarial work, cybersecurity, or socio-technical probing
- Native fluency in English and Dutch, both required
- Curious and adversarial mindset, with an instinct to push systems to breaking points
- Structured approach, using frameworks or benchmarks rather than ad hoc methods
- Clear communicator, able to explain risks to technical and non-technical stakeholders
- Adaptable, comfortable moving across projects and different customer contexts
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction
- Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
- Socio-technical risk expertise, for harassment, disinformation probing, or conversational AI abuse analysis
- Creative probing skills from psychology, acting, or writing to enable unconventional adversarial thinking
- Uncovering vulnerabilities that automated tests miss
- Delivering reproducible artifacts that strengthen customer AI systems
- Expanding evaluation coverage so fewer scenarios surprise production systems
- Increasing customer trust in their AI because adversarial vulnerabilities were identified and addressed
- Remote work only
- Hourly engagement, paid on an hourly basis
- All work is text-based
- Participation in higher-sensitivity projects is optional, topics will be communicated in advance, and wellness resources and clear guidelines are provided
48 - 62 hourly
Eligibility- Native fluency in both English and Dutch is required
- This role is remote; confirm you can perform duties from your location