Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Expert for English & Malay Red Teaming

$17–$25/hr

RemoteContracttechnology
Apply Now

About this role

Role Overview

Perform adversarial testing of conversational AI and agents to find real-world vulnerabilities and produce reproducible red team data that improves model safety. You will probe models with jailbreaks, prompt injections, and misuse scenarios, review outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors, and create artifacts customers can act on. All work is text-based. Participation in higher-sensitivity reviews is optional and supported by clear guidelines and wellness resources, and topics will be communicated before exposure.

Key Responsibilities
  • Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Review AI outputs related to sensitive topics and identify problematic behavior
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and comparable
  • Document reproducibly, producing reports, datasets, and attack cases that customers can act on
Qualifications
  • Native fluency in English and Malay, both languages required
  • Prior red teaming experience, for example AI adversarial work, cybersecurity, or socio-technical probing
  • Curious and adversarial mindset, with an instinct to push systems toward failure modes
  • Structured approach, comfortable using frameworks, benchmarks, or playbooks rather than ad hoc methods
  • Clear communicator, able to explain risks to both technical and non-technical stakeholders
  • Adaptable, able to move across projects and customers
Nice-to-Have Specialties
  • Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attack techniques, or model extraction
  • Cybersecurity skills such as penetration testing, exploit development, or reverse engineering
  • Socio-technical risk experience, for example harassment, disinformation probing, abuse analysis, or conversational AI testing
  • Creative probing skills, for example psychology, acting, or creative writing applied to adversarial scenarios
What Success Looks Like
  • You uncover vulnerabilities that automated tests miss
  • You deliver reproducible artifacts that help customers strengthen their AI systems
  • Evaluation coverage expands and fewer unexpected failures occur in production
  • Customers gain greater trust in their systems because adversarial scenarios have already been exercised
Work Terms

Remote, hourly engagement. All tasks are text-based. Participation in higher-sensitivity projects is optional and will be governed by clear guidelines and wellness resources. Topics that could be sensitive will be communicated to you before you are exposed to that content.

Compensation

17 - 25 hourly

Eligibility
  • Must have native fluency in both English and Malay
  • Remote work is required

Related Jobs