About this role
Role Overview
Join a red team focused on adversarial testing of conversational AI, probing models with adversarial inputs, surfacing vulnerabilities, and producing text-based red team data that helps customers make their systems safer. Work will involve reviewing model outputs that touch on sensitive areas such as bias, misinformation, and potentially harmful behaviors. Participation in higher-sensitivity reviews is optional, and you will receive advance notice of sensitive topics along with wellness resources and clear guidance.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Generate high-quality human data, annotate failures, classify vulnerabilities, and flag systemic risks
- Follow defined taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible
- Document findings clearly and reproducibly, producing reports, datasets, and attack cases customers can act on
- Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing
- A curious, adversarial mindset, with an instinct to push systems to breaking points
- A structured approach, using frameworks or benchmarks rather than ad hoc methods
- Strong communication skills, able to explain risks to technical and non-technical stakeholders
- Adaptability to move across projects and customers
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, and model extraction
- Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
- Socio-technical risk expertise, including harassment, disinformation probing, abuse analysis, and conversational AI testing
- Creative probing skills from psychology, acting, or writing to enable unconventional adversarial thinking
- Location: Remote
- Employment type: Hourly
- All work is text-based
- Participation in higher-sensitivity projects is optional, topics will be communicated before exposure, and wellness resources and clear guidelines are provided
- Pay range: 20 - 22 hourly
- Native fluency in English and Assamese is required