About this role
Role Overview
Join a red team that probes conversational AI to find vulnerabilities before they reach customers. This text-based role focuses on adversarial testing, documenting failures, and producing actionable datasets and reports that help make AI systems more robust and safe. Some projects will involve sensitive topics such as bias, misinformation, or harmful behaviors; participation in higher-sensitivity work is optional and supported with clear guidance and wellness resources. Topics are communicated in advance of exposure.
Key Responsibilities- Adversarially test conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible
- Document findings clearly, producing reports, reproducible datasets, and concrete attack cases customers can act on
- Prior red teaming experience, for example AI adversarial work, cybersecurity, or socio-technical probing
- Comfort with adversarial and curiosity-driven testing, pushing systems to identify breaking points
- Ability to work in structured ways, using frameworks or benchmarks rather than random approaches
- Strong communication skills, able to explain risks to both technical and non-technical stakeholders
- Adaptability to move across projects and customers
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attack techniques, and model extraction
- Cybersecurity experience, for example penetration testing, exploit development, or reverse engineering
- Socio-technical risk expertise, such as harassment or disinformation probing and abuse analysis for conversational AI
- Creative probing skills, including psychology, acting, or creative writing for unconventional adversarial scenarios
- Discovering vulnerabilities that automated tests do not detect
- Delivering reproducible artifacts that strengthen customer AI systems
- Expanding evaluation coverage so fewer surprising failures occur in production
- Increasing customer trust in the safety of their AI because systems have been probed like an adversary
- Location: Remote
- Employment type: Hourly engagement
- All tasks are text-based
- Participation in higher-sensitivity projects is optional and will include clear topic briefings and wellness resources
- Pay rate: 17 to 25 hourly
- Native fluency in both English and Indonesian is required
- Work performed remotely; candidates must be able to work remotely in accordance with any applicable local regulations