About this role
Role Overview
Probe conversational AI systems with adversarial inputs to identify vulnerabilities, generate high-quality safety data, and help make AI systems more robust and trustworthy. This text-based work includes reviewing model outputs that may address sensitive subjects such as bias, misinformation, and harmful behavior.
Key Responsibilities- Red team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Annotate failures, classify vulnerabilities, and flag systemic risks.
- Apply testing taxonomies, benchmarks, and playbooks to maintain consistent evaluations.
- Create reproducible reports, datasets, and attack cases that teams can use to strengthen AI systems.
- Expand evaluation coverage by testing more scenarios and uncovering vulnerabilities that automated tests can miss.
- Native fluency in both English and Dutch is required.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- A naturally adversarial, curious approach to testing systems and finding breaking points.
- Experience using structured frameworks or benchmarks rather than relying on unstructured testing.
- Ability to clearly communicate risks to technical and non-technical stakeholders.
- Adaptability to work across projects and customer contexts.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk evaluation, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
- Remote, hourly engagement.
- All work is text-based.
- Higher-sensitivity projects are optional and supported by clear guidelines and wellness resources. Topics will be communicated before any content exposure.
$48 to $62 per hour.