Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Red Teamer, English & Finnish

$48–$62/hr

RemoteContracttechnology
Apply Now

About this role

Help strengthen conversational AI by identifying vulnerabilities before they reach real-world users. In this remote, text-based role, you will adversarially test AI models and agents, generate actionable safety data, and document findings that help improve AI reliability and trustworthiness.

Key Responsibilities
  • Red team conversational AI models and agents through jailbreaks, prompt injection attempts, misuse cases, bias exploitation, and multi-turn manipulation.
  • Review AI outputs involving sensitive topics, including bias, misinformation, and harmful behaviors.
  • Annotate failures, classify vulnerabilities, and flag systemic risks to create high-quality human data.
  • Use established taxonomies, benchmarks, and playbooks to conduct consistent testing.
  • Create reproducible reports, datasets, and attack cases that technical teams can use to address identified issues.
Qualifications
  • Native fluency in both English and Finnish is required.
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • Ability to identify and test system weaknesses using structured frameworks or benchmarks.
  • Clear communication skills for explaining risks to technical and non-technical stakeholders.
  • Comfort working across changing projects and customer needs.
Preferred Experience
  • Adversarial machine learning, including jailbreak datasets, prompt injection, RLHF or DPO attacks, or model extraction.
  • Cybersecurity, including penetration testing, exploit development, or reverse engineering.
  • Socio-technical risk testing, such as harassment or disinformation probing, abuse analysis, or conversational AI testing.
  • Creative adversarial thinking informed by psychology, acting, or writing.
Work Terms
  • Remote, hourly engagement.
  • All work is text-based.
  • Higher-sensitivity assignments are optional and include clear guidelines and wellness resources. Topics will be communicated before any exposure to related content.
Compensation
  • $48 to $62 per hour.

Related Jobs

More like this