About this role
Role Overview
Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and creating high-quality red team data. This remote, text-based role focuses on uncovering risks that automated testing may miss and translating findings into actionable safety improvements.
Key Responsibilities
- Test conversational AI models and agents for jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation.
- Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors. Participation in higher-sensitivity projects is optional, with topics communicated in advance and supported by clear guidelines and wellness resources.
- Annotate failures, classify vulnerabilities, and flag systemic risks.
- Use taxonomies, benchmarks, and testing playbooks to conduct consistent evaluations.
- Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
Qualifications
- Native fluency in both English and Assamese is required.
- Prior experience in AI red teaming, cybersecurity, or socio-technical risk analysis.
- Ability to think adversarially, use structured testing frameworks, and clearly explain risks to technical and non-technical audiences.
- Comfort adapting across projects and customer needs.
Preferred Experience
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical safety evaluation, including harassment, misinformation, abuse analysis, or conversational AI testing.
- Creative adversarial approaches informed by psychology, acting, or writing.
Work Terms
- Remote hourly engagement.
Compensation
- $20 to $22 per hour.