About this role
Help strengthen conversational AI by probing models with adversarial inputs, identifying vulnerabilities, and producing actionable safety data. This remote, text-based role focuses on reviewing AI outputs, including content related to bias, misinformation, and harmful behavior. Participation in higher-sensitivity projects is optional, with topics communicated in advance and supported by clear guidelines and wellness resources.
Key Responsibilities- Red-team conversational AI models and agents through jailbreaks, prompt injection, misuse scenarios, bias exploitation, and multi-turn manipulation.
- Review and annotate AI failures, classify vulnerabilities, and flag systemic risks.
- Apply established taxonomies, benchmarks, and playbooks to maintain consistent testing.
- Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
- Native fluency in both English and Malay is required.
- Prior experience in AI red teaming, cybersecurity, or socio-technical risk probing.
- Ability to use structured frameworks or benchmarks and explain risks to both technical and non-technical audiences.
- Comfort working across varied projects and customer needs.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk analysis, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
- Remote, hourly engagement.
- $17 to $25 per hour.
- Uncover vulnerabilities that automated testing misses.
- Deliver reproducible artifacts that improve AI-system safety.
- Expand evaluation coverage across more scenarios and reduce production surprises.