About this role
Role Overview
Lead adversarial testing of conversational AI systems to find and document vulnerabilities related to bias, misinformation, and other harmful behaviors. You will design and execute text-based attacks, produce reproducible human-data artifacts, and follow structured playbooks so engineering and safety teams can remediate issues.
Key Responsibilities- Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Review AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors, and clearly flag risky or problematic responses.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and surfacing systemic risks.
- Follow taxonomies, benchmarks, and playbooks to keep testing consistent and comparable across projects.
- Produce reproducible deliverables, including reports, datasets, and documented attack cases that engineers and customers can act on.
- Optionally participate in higher-sensitivity projects, when available, under clear content guidelines and with access to wellness resources.
- Native fluency in both English and Thai, required.
- Prior red teaming experience, such as adversarial AI testing, cybersecurity offensive work, or socio-technical probing.
- A curious and adversarial mindset, with a tendency to probe systems to breaking points.
- Structured approach to testing, using frameworks or benchmarks rather than ad hoc methods.
- Clear communicator, able to explain risks to technical and non-technical stakeholders.
- Adaptable, comfortable moving across projects and customers.
- Adversarial machine learning, for example jailbreak datasets, prompt injection, RLHF/DPO attack techniques, or model extraction.
- Cybersecurity skills such as penetration testing, exploit development, or reverse engineering.
- Socio-technical risk experience, for example harassment or disinformation probing, abuse analysis, or conversational AI testing focused on real-world harms.
- Creative probing abilities, including psychology, acting, or writing for unconventional adversarial scenarios.
- Uncovered vulnerabilities that automated tests miss.
- Deliverables that are reproducible and actionable for engineering teams and customers.
- Expanded evaluation coverage across scenarios, resulting in fewer surprises in production.
- Greater customer confidence in their AI safety because adversarial testing has exposed and helped mitigate risks.
- Location: Remote, text-based work.
- Employment type: Hourly engagement.
- Assignments are project-based and will vary by project and customer.
- Participation in higher-sensitivity projects is optional, topics will be communicated before exposure, and wellness resources and clear guidelines are provided when sensitive content is involved.
Pay range: $24 to $35 hourly.
Eligibility- Native fluency in English and Thai is required.
- Remote work is required; candidates must be able to perform all duties remotely.