About this role
Help strengthen conversational AI by identifying vulnerabilities before they reach real-world users. In this remote, text-based role, you will adversarially test AI models and agents, generate actionable safety data, and document findings that help improve AI reliability and trustworthiness.
Key Responsibilities- Red team conversational AI models and agents through jailbreaks, prompt injection attempts, misuse cases, bias exploitation, and multi-turn manipulation.
- Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors.
- Annotate failures, classify vulnerabilities, and flag systemic risks to create high-quality human data.
- Use established taxonomies, benchmarks, and playbooks to conduct consistent testing.
- Create reproducible reports, datasets, and attack cases that technical teams can use to address identified issues.
- Native fluency in both English and Finnish is required.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to identify and test system weaknesses using structured frameworks or benchmarks.
- Clear communication skills for explaining risks to technical and non-technical stakeholders.
- Comfort working across changing projects and customer needs.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk testing, such as harassment or disinformation probing, abuse analysis, or conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
- Remote, hourly engagement.
- All work is text-based.
- Higher-sensitivity assignments are optional and include clear guidelines and wellness resources. Topics will be communicated before any exposure to related content.
- $48 to $62 per hour.