About this role
Probe conversational AI systems for weaknesses before they reach users. In this remote, text-based role, you will use adversarial testing to uncover vulnerabilities, create high-quality safety data, and produce practical findings that help strengthen AI systems.
Role OverviewYou will review AI outputs that may address sensitive subjects, including bias, misinformation, and harmful behaviors. Participation in projects involving higher-sensitivity content is optional; topics are disclosed before exposure, and clear guidelines and wellness resources are provided.
Key Responsibilities- Red team conversational AI models and agents through jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Annotate failures, classify vulnerabilities, and identify systemic risks.
- Apply established taxonomies, benchmarks, and playbooks to conduct consistent testing.
- Create reproducible reports, datasets, and attack cases that teams can use to address identified risks.
- Help expand evaluation coverage by testing more scenarios and surfacing issues that automated tests may miss.
- Native fluency in both English and Swedish is required.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to think adversarially, test systems at their limits, and work with structured frameworks or benchmarks.
- Clear communication skills for explaining risks to both technical and non-technical stakeholders.
- Comfort moving across projects and working with different customer needs.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk assessment, including harassment or disinformation probing, abuse analysis, or conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
- Remote hourly engagement.
- $48 to $62 per hour.