About this role
Role Overview
Work as a red team human data expert attacking conversational AI in English and Dutch to surface vulnerabilities, produce reproducible adversarial test cases, and generate the human-labeled data customers use to improve model safety. The role focuses on text-only interactions and on probing areas such as bias, misinformation, and harmful behavior.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structure and consistency by following taxonomies, benchmarks, and playbooks during testing.
- Document findings reproducibly, producing reports, datasets, and attack cases customers can act on.
- Work with content that may touch sensitive topics; all tasks are text-based, higher-sensitivity projects are optional, and you will be given clear guidance and access to wellness resources. Topics will be communicated before exposure.
- Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing.
- Comfort with adversarial thinking, pushing systems toward failure modes, and explaining risks clearly to both technical and non-technical stakeholders.
- Structured approach to testing, using frameworks or benchmarks rather than ad hoc methods.
- Adaptability to move across projects and customers while maintaining consistent documentation and outputs.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attack experience, and model extraction.
- Cybersecurity skills such as penetration testing, exploit development, or reverse engineering.
- Socio-technical risk expertise, for example harassment, disinformation probing, abuse analysis, or conversational AI testing.
- Creative probing skills, including psychology, acting, or writing for unconventional adversarial approaches.
- Uncovering vulnerabilities that automated tests miss.
- Delivering reproducible artifacts that help strengthen customer AI systems.
- Expanding evaluation coverage so more scenarios are tested and fewer surprises occur in production.
- Enabling customer trust in their AI by proactively finding and documenting adversarial failures.
- Location: Remote.
- Employment type: hourly engagement.
- All tasks are text-based. Participation in higher-sensitivity projects is optional, with pre-notification of sensitive topics and access to clear guidelines and wellness support.
- Hourly rate: 48 - 62 hourly.
- Native fluency in both English and Dutch is required for this role.