About this role
Role Overview
Help strengthen conversational AI systems by probing them with adversarial inputs, identifying weaknesses, and producing human-generated safety data. This text-based role focuses on reviewing AI outputs, including content involving bias, misinformation, or harmful behaviors. Participation in higher-sensitivity work is optional; topics are communicated in advance, with clear guidelines and wellness resources available.
Key Responsibilities
- Test conversational AI models and agents for jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation.
- Annotate failures, classify vulnerabilities, and flag systemic risks.
- Apply testing taxonomies, benchmarks, and playbooks consistently.
- Create reproducible reports, datasets, and attack cases that teams can use to improve AI systems.
- Expand evaluation coverage by testing more scenarios and surfacing vulnerabilities that automated testing may miss.
Qualifications
- Native fluency in both English and Bengali is required.
- Prior red-teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Curiosity and an adversarial mindset for pushing systems to their limits.
- Ability to use structured frameworks or benchmarks and communicate risks to technical and non-technical stakeholders.
- Adaptability across projects and customers.
Preferred Specialties
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity, including penetration testing, exploit development, or reverse engineering.
- Socio-technical risk assessment, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Creative probing informed by psychology, acting, or writing.
Work Terms
- Remote, hourly engagement.
- All work is text-based.
Compensation
$20 to $22 per hour.
Eligibility
This role requires native-level English and Bengali language skills.