Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Red Teamer, English & Dutch

$48–$62/hr

RemoteContracttechnology
Apply Now

About this role

Role Overview

Probe conversational AI systems with adversarial inputs to identify vulnerabilities, generate high-quality safety data, and help make AI systems more robust and trustworthy. This text-based work includes reviewing model outputs that may address sensitive subjects such as bias, misinformation, and harmful behavior.

Key Responsibilities
  • Red team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
  • Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply testing taxonomies, benchmarks, and playbooks to maintain consistent evaluations.
  • Create reproducible reports, datasets, and attack cases that teams can use to strengthen AI systems.
  • Expand evaluation coverage by testing more scenarios and uncovering vulnerabilities that automated tests can miss.
Qualifications
  • Native fluency in both English and Dutch is required.
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • A naturally adversarial, curious approach to testing systems and finding breaking points.
  • Experience using structured frameworks or benchmarks rather than relying on unstructured testing.
  • Ability to clearly communicate risks to technical and non-technical stakeholders.
  • Adaptability to work across projects and customer contexts.
Preferred Specialties
  • Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
  • Cybersecurity, including penetration testing, exploit development, or reverse engineering.
  • Socio-technical risk evaluation, including harassment or misinformation probing, abuse analysis, or conversational AI testing.
  • Creative adversarial thinking informed by psychology, acting, or writing.
Work Terms
  • Remote, hourly engagement.
  • All work is text-based.
  • Higher-sensitivity projects are optional and supported by clear guidelines and wellness resources. Topics will be communicated before any content exposure.
Compensation

$48 to $62 per hour.

Related Jobs

More like this