Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Expert for English and Finnish Red Teaming

$48–$62/hr

RemoteContracttechnology
Apply Now

About this role

Role Overview

Probe conversational AI systems as an adversary to uncover vulnerabilities that automated checks miss. You will design and execute red team tests, review model outputs that touch on sensitive topics such as bias, misinformation, and harmful behaviors, and deliver reproducible attack cases, labeled datasets, and reports customers can act on. All work is text-based, and participation in higher-sensitivity tasks is optional with clear guidance and wellness support. Topics will be communicated before exposure.

Key Responsibilities
  • Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks
  • Apply structure by following taxonomies, benchmarks, and playbooks to keep testing consistent
  • Document reproducibly by producing reports, datasets, and attack cases customers can act on
  • Review AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors
Qualifications
  • Native fluency in English and Finnish is required
  • Prior red teaming experience, such as AI adversarial work, cybersecurity, or socio-technical probing
  • A curious, adversarial mindset that instinctively pushes systems to breaking points
  • Structured approach, using frameworks or benchmarks rather than unguided attempts
  • Strong communication skills, able to explain risks to technical and non-technical stakeholders
  • Adaptability to move across projects and customers
Nice-to-Have Specialties
  • Adversarial machine learning: jailbreak datasets, prompt injection, RLHF or DPO attacks, model extraction
  • Cybersecurity: penetration testing, exploit development, reverse engineering
  • Socio-technical risk: harassment and disinformation probing, abuse analysis, conversational AI testing
  • Creative probing skills: psychology, acting, or creative writing for unconventional adversarial approaches
What Success Looks Like
  • Uncovering vulnerabilities that automated tests miss
  • Delivering reproducible artifacts that materially strengthen customer AI systems
  • Expanding evaluation coverage so fewer surprising behaviors reach production
  • Increasing customer trust because models have been probed like real adversaries
Why Join

Work on human data-driven AI red teaming at the frontier of safety, and play a direct role in making AI systems more robust, safe, and trustworthy.

Work Terms
  • Remote engagement
  • Hourly work arrangement
  • All tasks are text-based
  • Higher-sensitivity projects are optional and accompanied by clear guidelines and wellness resources; topics are disclosed in advance
Compensation

48 - 62 hourly

Eligibility
  • Must be able to work remotely
  • Native fluency in both English and Finnish is required

Related Jobs