About this role
You will probe conversational AI models and agents with adversarial inputs to surface vulnerabilities, produce high-quality human red team data, and create reproducible attack cases and reports that help clients make their AI safer. All work is text-based and may involve outputs touching on sensitive topics such as bias, misinformation, or harmful behaviors.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
- Review AI outputs that relate to sensitive topics, documenting failures and systemic risks
- Generate high-quality human data, annotate failures, classify vulnerabilities, and flag patterns that automated tests miss
- Apply structure to testing by following taxonomies, benchmarks, and playbooks to ensure consistent coverage
- Document findings reproducibly, producing reports, datasets, and attack cases clients can act on
- Opt in, when comfortable, to higher-sensitivity projects that are governed by clear guidelines and supported with wellness resources; topics will be clearly communicated before exposure
- Native-level fluency in English and Portuguese, global variant, excluding Brazilian Portuguese, is required
- Prior red teaming experience, for example AI adversarial work, cybersecurity, or socio-technical probing
- A curious, adversarial mindset, with an instinct to push systems to breaking points
- Structured approach to testing, using frameworks or benchmarks rather than ad hoc methods
- Clear communicator, able to explain risks to both technical and non-technical stakeholders
- Adaptable, able to move across projects and clients
- Adversarial machine learning, such as jailbreak datasets, prompt injection, RLHF or DPO attack methods, and model extraction
- Cybersecurity skills, including penetration testing, exploit development, or reverse engineering
- Socio-technical risk experience, including harassment or disinformation probing and conversational AI abuse analysis
- Creative probing skills, such as psychology, acting, or creative writing for unconventional adversarial thinking
- You uncover vulnerabilities that automated tests do not detect
- You deliver reproducible artifacts that strengthen client AI systems
- Evaluation coverage expands, producing fewer surprises in production
- Clients trust the safety of their AI because it has been probed thoroughly by your work
Remote engagement, text-based work. Employment type is hourly. Participation in higher-sensitivity assignments is optional and accompanied by clear guidance and wellness resources. Topics that could be sensitive will be communicated to you before you are exposed to them.
CompensationPay range: 29 - 45 hourly.
Eligibility- This role requires native fluency in English and Portuguese, global variant, excluding Brazilian Portuguese
- Remote work is permitted