About this role
Role Overview
Probe conversational AI systems for safety and misuse vulnerabilities in both English and Norwegian, producing high-quality adversarial test cases, annotations, and reproducible reports that help customers reduce bias, misinformation, and harmful behaviors. The work is text-based and focused on uncovering failures automated tests miss.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Review AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors
- Generate human-labeled data: annotate failures, classify vulnerabilities, and flag systemic risks
- Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible
- Document findings as actionable artifacts, including reports, datasets, and reproducible attack cases customers can act on
- Native fluency in English and Norwegian is required
- Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing
- Comfort with structured testing methods, using frameworks or benchmarks rather than ad hoc approaches
- Strong communication skills, able to explain risks to technical and non-technical stakeholders
- Ability to adapt and move across projects and customer contexts
- Adversarial machine learning: jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction
- Cybersecurity: penetration testing, exploit development, reverse engineering
- Socio-technical risk analysis: harassment or disinformation probing, abuse analysis, conversational AI testing
- Creative probing skills: psychology, acting, or writing for unconventional adversarial thinking
- Discovering vulnerabilities automated tests miss
- Delivering reproducible artifacts that customers can use to strengthen their AI systems
- Expanding evaluation coverage so fewer surprise failure modes appear in production
- Increasing customer trust in the safety of their AI systems because adversarial scenarios have already been exercised
- Remote engagement
- Hourly employment
- All tasks are text-based
- Participation in higher-sensitivity projects is optional, with clear topic briefings and wellness resources provided; topics will be communicated before exposure
48 - 62 hourly
Eligibility- Native fluency in both English and Norwegian is required
- Ability to work remotely under an hourly engagement model