About this role
Role Overview
Probe conversational AI systems with adversarial inputs to uncover vulnerabilities, generate actionable safety data, and help strengthen AI systems before issues reach production. This text-based role includes reviewing AI outputs that may address sensitive topics, including bias, misinformation, and harmful behavior.
Key Responsibilities- Test conversational AI models and agents for jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Review sensitive AI outputs and identify failures, vulnerabilities, and systemic risks.
- Annotate findings, classify vulnerabilities, and create high-quality human data for AI safety work.
- Use taxonomies, benchmarks, and testing playbooks to conduct consistent evaluations.
- Produce reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
- Native fluency in both English and Thai is required.
- Prior experience with AI red teaming, cybersecurity, or socio-technical risk assessment.
- Ability to test systems methodically using frameworks or benchmarks and communicate risks to technical and non-technical audiences.
- Comfort adapting across projects and customer needs.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk analysis, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Psychology, acting, or writing skills that support unconventional adversarial thinking.
- Remote, hourly engagement.
- All work is text-based.
- Participation in higher-sensitivity projects is optional. Topics will be communicated before exposure, with clear guidelines and wellness resources available.
$24 to $35 per hour.