About this role
Role Overview
This role performs adversarial testing of conversational AI, producing reproducible human red team data that reveals model vulnerabilities and helps make AI safer for customers. Work is fully text based and focuses on outputs that touch on sensitive topics such as bias, misinformation, and harmful behaviors. Higher-sensitivity work is optional and supported with clear guidelines and wellness resources, and topics will be communicated before any exposure.
Key Responsibilities- Execute red team attacks on conversational models and agents, including jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation
- Review AI outputs that involve sensitive content and identify failure modes related to bias, misinformation, or harmful behavior
- Generate high quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Follow established taxonomies, benchmarks, and playbooks to maintain consistent testing methodology
- Document findings reproducibly, delivering reports, datasets, and attack cases that customers can act on
- Native fluency in English and Norwegian is required
- Prior red teaming or adversarial testing experience, such as AI adversarial work, cybersecurity, or socio-technical probing
- Comfort with adversarial mindset, pushing systems to their limits while following structured frameworks
- Strong written communication, able to explain technical and non-technical risks clearly
- Adaptability to move across projects and customers
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attack knowledge, model extraction
- Cybersecurity experience, such as penetration testing, exploit development, or reverse engineering
- Socio-technical risk expertise, including harassment, disinformation, abuse analysis, and conversational AI testing
- Creative probing skills, for example psychology, acting, or adversarial writing
- Location: Remote
- All work is text based
- Participation in higher-sensitivity projects is optional, with clear topic disclosure and wellness resources provided before exposure
Hourly rate, 48 - 62 hourly
Eligibility- Native fluency in both English and Norwegian is mandatory