About this role
Role Overview
Act as a human red teamer who probes conversational AI systems in English and Swedish to uncover safety vulnerabilities, produce high-quality adversarial inputs, and deliver reproducible artifacts customers can act on. The work is fully text-based, focuses on sensitive topics such as bias, misinformation, and harmful behaviors, and includes optional higher-sensitivity engagements supported by clear guidelines and wellness resources.
Key Responsibilities- Perform adversarial testing of conversational AI models and agents, including jailbreaks, prompt injection, misuse scenarios, bias exploitation, and multi-turn manipulation.
- Generate human-derived safety data: annotate model failures, classify types of vulnerabilities, and flag systemic risks.
- Follow established structure: use taxonomies, benchmarks, and playbooks to ensure consistent, comparable testing.
- Document findings reproducibly: produce reports, datasets, and attack cases that customers can reproduce and act on.
- Optionally participate in higher-sensitivity reviews, with topics communicated in advance and access to wellness resources when applicable.
- Native fluency in both English and Swedish, spoken and written, is required.
- Prior red teaming experience, such as AI adversarial work, cybersecurity testing, or socio-technical probing.
- A curious, adversarial mindset with a habit of pushing systems to breaking points while working methodically.
- Ability to communicate risks clearly to both technical and non-technical stakeholders.
- Comfort working across projects and adapting to different customers and requirements.
Preferred specialties
- Adversarial ML techniques, for example jailbreak datasets, prompt injection, RLHF or DPO attacks, and model extraction.
- Cybersecurity experience, such as penetration testing, exploit development, or reverse engineering.
- Socio-technical risk work, including harassment, disinformation probing, abuse analysis, and conversational AI testing.
- Creative probing skills from psychology, acting, or writing to design unconventional adversarial approaches.
- Remote engagement, text-only work.
- This is an hourly engagement, paid at the rates listed below.
- Participation in higher-sensitivity projects is optional, topics are disclosed before exposure, and wellness guidance and other support are provided for sensitive work.
- 48 - 62 hourly
- Native fluency in English and Swedish is mandatory.
- Ability to work remotely and to review AI outputs that may involve bias, misinformation, or harmful behaviors. Participation in sensitive-content reviews is voluntary and will be clearly communicated in advance.