Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Expert for English and Danish Red Teaming

$48–$62/hr

RemoteContracttechnology
Apply Now

About this role

Role Overview

Probe conversational AI systems with adversarial inputs to expose vulnerabilities, produce reproducible red-team data, and deliver actionable reports that help customers reduce bias, misinformation, and harmful behaviors. This is a remote, text-based role focused on rigorous, structured safety testing across multiple projects and customers.

Key Responsibilities
  • Red team conversational AI models and agents by developing jailbreaks, prompt injections, misuse scenarios, bias exploitation techniques, and multi-turn manipulation strategies
  • Review AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors, and generate human-reviewed examples of failures
  • Produce high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible
  • Document findings clearly and reproducibly, delivering reports, datasets, and attack cases customers can act on
  • Participation in higher-sensitivity projects is optional and will be supported by clear guidelines and wellness resources; topics will be clearly communicated before any exposure
Qualifications
  • Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing
  • A curious, adversarial mindset, with an instinct to push systems toward breaking points
  • Structured approach to testing, using frameworks, benchmarks, or playbooks rather than ad hoc methods
  • Strong communication skills, able to explain risks to both technical and non-technical stakeholders
  • Adaptability, able to move across projects and customer contexts

Nice-to-Have Specialties

  • Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attack knowledge, and model extraction techniques
  • Cybersecurity experience such as penetration testing, exploit development, or reverse engineering
  • Socio-technical risk expertise, for example harassment or disinformation probing, abuse analysis, and conversational AI testing
  • Creative probing skills from psychology, acting, or creative writing to design unconventional adversarial approaches
Work Terms
  • Remote, text-based work only
  • Employment type, hourly
  • Work may involve content on sensitive topics; participation in higher-sensitivity work is optional and will include clear guidance and access to wellness resources
  • Topics that could be sensitive will be communicated in advance of any exposure
Compensation

Pay rate: 48 - 62 hourly

Eligibility
  • Native fluency in both English and Danish is required
  • Role is remote; candidates must be able to perform all duties remotely

Related Jobs