About this role
Role Overview
Perform adversarial red teaming of conversational AI to find real vulnerabilities, generate high-quality human data, and deliver reproducible attack cases that help customers harden their systems. The work is fully text based and may involve reviewing outputs that touch on sensitive topics, including bias, misinformation, and harmful behaviors. Participation in higher-sensitivity tasks is optional and will be supported with clear guidelines and wellness resources, with topics communicated to you before exposure.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Follow and apply taxonomies, benchmarks, and playbooks to keep testing consistent and structured
- Document findings reproducibly, producing reports, datasets, and attack cases that customers can act on
- Prior red teaming experience, for example AI adversarial work, cybersecurity, or socio-technical probing
- Native fluency in English and Vietnamese is required
- Curious and adversarial mindset, with a tendency to push systems to breaking points
- Comfort with structured approaches, using frameworks or benchmarks rather than ad hoc methods
- Strong communication skills, able to explain risks to both technical and non-technical stakeholders
- Adaptable to shifting projects and customer requirements
- Adversarial ML, including jailbreak datasets, prompt injection, RLHF or DPO attack experience, model extraction
- Cybersecurity skills such as penetration testing, exploit development, or reverse engineering
- Socio-technical risk assessment experience, for example harassment, disinformation probing, or conversational abuse analysis
- Creative probing skills, including psychology, acting, or writing for unconventional adversarial thinking
- Uncovering vulnerabilities automated tests miss
- Delivering reproducible artifacts that directly strengthen customer AI systems
- Expanding evaluation coverage so more scenarios are tested and fewer surprises occur in production
- Customers gain measurable trust in their AI because adversarial issues were proactively uncovered and addressed
- Remote work
- Hourly engagement
- All tasks are text based
- Participation in any higher-sensitivity projects is optional, will be preceded by topic disclosure, and will include clear guidelines and wellness support
- Pay range: 17 - 25 hourly
- Native fluency in both English and Vietnamese is required
- Remote work is required and expected