About this role
Role Overview
Probe conversational AI systems with adversarial inputs to expose vulnerabilities, produce reproducible red-team data, and deliver actionable reports that help customers reduce bias, misinformation, and harmful behaviors. This is a remote, text-based role focused on rigorous, structured safety testing across multiple projects and customers.
Key Responsibilities- Red team conversational AI models and agents by developing jailbreaks, prompt injections, misuse scenarios, bias exploitation techniques, and multi-turn manipulation strategies
- Review AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors, and generate human-reviewed examples of failures
- Produce high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible
- Document findings clearly and reproducibly, delivering reports, datasets, and attack cases customers can act on
- Participation in higher-sensitivity projects is optional and will be supported by clear guidelines and wellness resources; topics will be clearly communicated before any exposure
- Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing
- A curious, adversarial mindset, with an instinct to push systems toward breaking points
- Structured approach to testing, using frameworks, benchmarks, or playbooks rather than ad hoc methods
- Strong communication skills, able to explain risks to both technical and non-technical stakeholders
- Adaptability, able to move across projects and customer contexts
Nice-to-Have Specialties
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attack knowledge, and model extraction techniques
- Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
- Socio-technical risk expertise, for example harassment or disinformation probing, abuse analysis, and conversational AI testing
- Creative probing skills from psychology, acting, or creative writing to design unconventional adversarial approaches
- Remote, text-based work only
- Employment type, hourly
- Work may involve content on sensitive topics; participation in higher-sensitivity work is optional and will include clear guidance and access to wellness resources
- Topics that could be sensitive will be communicated in advance of any exposure
Pay rate: 48 - 62 hourly
Eligibility- Native fluency in both English and Danish is required
- Role is remote; candidates must be able to perform all duties remotely