About this role
Probe conversational AI systems with adversarial inputs to reveal vulnerabilities, produce reproducible red team data, and help customers remediate risks. The work focuses on text-based review of outputs that may touch on sensitive topics such as bias, misinformation, or harmful behaviors. Participation in higher-sensitivity projects is optional and supported with clear guidelines and wellness resources, and topics will be clearly communicated before any exposure.
Key Responsibilities- Perform red teaming of conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Review AI outputs that involve sensitive topics, identify problematic behaviors, and surface patterns of risk
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
- Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible
- Document findings in actionable artifacts, including reports, datasets, and reproducible attack cases for customers
- Native fluency in English and Portuguese, global variant, excluding Brazilian Portuguese, is required
- Prior red teaming or adversarial experience, such as AI adversarial work, cybersecurity, or socio-technical probing
- Natural curiosity and an adversarial mindset, with a drive to push systems to breaking points
- Structured approach, using frameworks or benchmarks rather than ad hoc testing
- Strong communication skills, able to explain risks to both technical and non-technical stakeholders
- Adaptability to move across projects and customers
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO related attacks, and model extraction
- Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
- Socio-technical risk expertise, including harassment and disinformation probing, abuse analysis, and conversational AI testing
- Creative probing skills from psychology, acting, or creative writing for unconventional adversarial thinking
- Uncovering vulnerabilities that automated tests miss
- Delivering reproducible artifacts that help strengthen customer AI systems
- Expanding evaluation coverage so more scenarios are tested and there are fewer surprises in production
- Helping customers trust their AI because systems have been probed from an adversarial perspective
Build experience in human data driven AI red teaming at the frontier of safety, and play a direct role in making AI systems more robust, safe, and trustworthy.
Work Terms- Remote role, text based work
- Employment type, hourly
- Participation in higher-sensitivity content projects is optional, with clear guidelines and wellness resources provided
- Topics that may be sensitive will be communicated in advance of exposure
29 - 45 hourly
Eligibility- Native fluency in English and Portuguese, global variant, excluding Brazilian Portuguese, is required
- Remote candidates accepted