About this role
Role Overview
Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and creating actionable safety data. This text-based remote role includes reviewing AI outputs involving sensitive subjects, including bias, misinformation, and harmful behavior. Higher-sensitivity assignments are optional, supported by clear guidelines and wellness resources, and content topics will be communicated before exposure.
Key Responsibilities
- Red team conversational AI models and agents through jailbreaks, prompt injection, misuse scenarios, bias exploitation, and multi-turn manipulation.
- Annotate model failures, classify vulnerabilities, and flag systemic risks.
- Use established taxonomies, benchmarks, and playbooks to conduct consistent testing.
- Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
- Expand evaluation coverage by identifying issues automated testing may miss.
Qualifications
- Native-level fluency in both English and Indonesian is required.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to apply structured frameworks or benchmarks, communicate risks to technical and non-technical audiences, and move between projects and customers.
Preferred Experience
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk assessment, including harassment or disinformation probing, abuse analysis, or conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
Work Terms
- Remote, hourly engagement.
Compensation
- $17 to $25 per hour.