About this role
Role Overview
Help strengthen conversational AI systems by probing them with adversarial inputs, identifying vulnerabilities, and producing actionable safety data. This remote, text-based role focuses on uncovering risks that automated testing can miss, including bias, misinformation, harmful behavior, misuse, and manipulation.
Key Responsibilities
- Test conversational AI models and agents for jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation.
- Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors.
- Annotate model failures, classify vulnerabilities, and flag systemic risks.
- Follow established taxonomies, benchmarks, and playbooks to conduct consistent testing.
- Create reproducible reports, datasets, and attack cases that stakeholders can use to improve AI systems.
- Expand evaluation coverage across scenarios and help reduce unexpected production risks.
Qualifications
- Native fluency in both English and Punjabi is required.
- Prior red-teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
- Ability to think adversarially, apply structured testing frameworks, and communicate risks clearly to technical and non-technical audiences.
- Comfort adapting across projects and customer environments.
Preferred Experience
- Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
- Cybersecurity experience such as penetration testing, exploit development, or reverse engineering.
- Socio-technical risk work, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
- Creative adversarial thinking informed by psychology, acting, or writing.
Work Terms
- Remote, hourly engagement.
- All work is text-based.
- Participation in higher-sensitivity projects is optional. Topics will be communicated before exposure, with clear guidelines and wellness resources available.
Compensation
$20 to $22 per hour.