Key details
- Role type
- Contract
- Compensation
- $17–$25/hr
- Work arrangement
- Remote
- Category
- technology
- Confirmed requirements
- 4
About this role
Help strengthen conversational AI by testing it from an adversarial perspective. In this remote role, you will probe models and agents for weaknesses, create actionable safety data, and help identify risks before they reach production.
Role OverviewThis text-based work includes reviewing AI outputs that may involve sensitive subjects, including bias, misinformation, and harmful behavior. Participation in higher-sensitivity projects is optional, with clear guidance and wellness resources. Topics will be communicated before any exposure.
Key Responsibilities- Test conversational AI models and agents for jailbreaks, prompt injection, misuse, bias exploitation, and multi-turn manipulation.
- Annotate model failures, classify vulnerabilities, and flag systemic risks.
- Use taxonomies, benchmarks, and playbooks to conduct consistent testing.
- Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
- Expand evaluation coverage by identifying vulnerabilities that automated testing may miss.
- Native fluency in both English and Indonesian is required.
- Prior experience in AI red teaming, adversarial AI work, cybersecurity, or socio-technical probing.
- Ability to apply structured frameworks and benchmarks to adversarial testing.
- Clear communication skills for explaining risks to technical and non-technical audiences.
- Comfort moving across projects and customer contexts.
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity work such as penetration testing, exploit development, or reverse engineering.
- Socio-technical risk analysis, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Creative disciplines such as psychology, acting, or writing that support unconventional adversarial thinking.
- Remote, hourly engagement.
- $17 to $25 per hour.
What to prepare before applying
- Remote work location.
- Native fluency in English and Indonesian.
- Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Adversarial mindset for probing systems to breaking points.
These are the confirmed hard requirements. The Apply button routes you to the partner platform where you complete the application.