About this role
Role Overview
This role probes conversational AI models to find real-world safety weaknesses and produce reproducible red team data that customers can act on. You will run adversarial tests, document failures, and help improve model robustness. All work is text based and may involve outputs touching sensitive topics such as bias, misinformation, or harmful behaviors. Participation in higher-sensitivity reviews is optional and will be supported by clear guidelines and wellness resources. Topics will be communicated to you before any exposure.
Key Responsibilities- Conduct red teaming of conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Apply structure and consistency by following taxonomies, benchmarks, and playbooks during testing
- Document findings reproducibly, producing reports, datasets, and attack cases that customers can reproduce and act on
- Native fluency in English and Finnish is required
- Prior red teaming experience, such as AI adversarial testing, cybersecurity, or socio-technical probing
- Comfort with adversarial thinking, curiosity, and a tendency to push systems to their limits
- Ability to work in a structured way using frameworks or benchmarks rather than random testing
- Clear communication skills, able to explain technical and non-technical risks
- Adaptability to move between different projects and customer contexts
Nice-to-have specialties
- Adversarial machine learning: jailbreak dataset creation, prompt injection, RLHF or DPO attack techniques, model extraction
- Cybersecurity skills: penetration testing, exploit development, reverse engineering
- Socio-technical risk expertise: harassment and disinformation probing, abuse analysis, conversational AI safety testing
- Creative probing skills: psychology, acting, or writing that enable unconventional adversarial approaches
- Location: Remote
- Employment type: hourly
- All tasks are text based
- Participation in higher-sensitivity projects is optional and accompanied by guidance and wellness resources
Rate: 48 - 62 hourly
Eligibility- Native fluency in English and Finnish is required
- Candidates must be able to work remotely
- This project involves reviewing outputs that may cover sensitive subject matter such as bias, misinformation, or harmful behavior; topics will be disclosed before exposure