About this role
Role Overview
Help strengthen conversational AI by playing the role of an adversary. This position focuses on red teaming language models and agents, producing reproducible attack cases, annotated datasets, and clear vulnerability reports that customers can act on. Work is text based and centers on finding failures related to bias, misinformation, harmful behavior, and other safety risks.
Key Responsibilities- Conduct adversarial testing of conversational AI systems, including jailbreaks, prompt injection, misuse cases, bias exploitation, and multi turn manipulation.
- Generate high quality human data by annotating model failures, classifying vulnerabilities, and flagging systemic risks.
- Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and repeatable.
- Document findings in reproducible form, producing reports, datasets, and attack cases that customers can reproduce and remediate.
- Review AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors, with optional participation in higher sensitivity projects.
- Native fluency in English and Swedish is required.
- Prior red teaming experience, including AI adversarial work, cybersecurity testing, or socio technical probing.
- Ability to think adversarially and push systems toward breaking points while staying structured and methodical.
- Skilled at explaining risks clearly to both technical and non technical stakeholders.
- Comfortable adapting across projects and customers, working in changing scopes and contexts.
- Adversarial machine learning, such as jailbreak datasets, prompt injection, RLHF or DPO attack methods, and model extraction.
- Cybersecurity experience, including penetration testing, exploit development, or reverse engineering.
- Socio technical risk expertise, for example harassment, disinformation probing, abuse analysis, or conversational AI testing.
- Creative probing skills from psychology, acting, or writing for unconventional adversarial thinking.
- Discovering vulnerabilities that automated tests miss.
- Delivering reproducible artifacts that customers use to improve AI safety.
- Expanding evaluation coverage so more scenarios are tested and fewer surprises occur in production.
- Increasing customer trust in deployed AI systems by proactively finding and documenting risks.
- Location, remote only.
- All work is text based.
- Participation in higher sensitivity projects is optional, and those projects are accompanied by clear guidelines and wellness resources.
- Topics that may contain sensitive content will be clearly communicated before exposure.
- Employment type, hourly.
48 - 62 hourly
Eligibility- Native fluency in English and Swedish is required for this position.