Skip to content
SaidGig
Sign up.

Give me your email, I promise I won't do anything weird with it.

AI Safety Expert for English and Vietnamese

$17–$25/hr

RemoteContracttechnology
Apply Now

About this role

Role Overview

Probe conversational AI systems to find vulnerabilities and produce reproducible adversarial data that customers can use to harden their models. This is a text-based red teaming role focused on exposing bias, misinformation, harmful behaviors, jailbreaks, and other misuse cases, then documenting findings in structured, actionable artifacts. Participation in higher-sensitivity reviews is optional and supported with clear guidelines and wellness resources, and topics will be communicated before any exposure.

Key Responsibilities
  • Red team conversational AI models and agents by designing and executing adversarial interactions, including jailbreaks, prompt injection, misuse cases, bias exploitation, and multi-turn manipulation.
  • Generate high-quality human data: annotate model failures, classify vulnerabilities, and flag systemic risks.
  • Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and reproducible.
  • Document results clearly, delivering reports, datasets, and attack case examples customers can act on.
Qualifications
  • Prior red teaming experience, such as AI adversarial work, cybersecurity testing, or socio-technical probing.
  • Native fluency in English and Vietnamese, with strong written and verbal communication in both languages.
  • Curious and adversarial mindset, with an instinct to push systems to failure points rather than accepting surface behavior.
  • Structured approach to testing, using frameworks and benchmarks rather than ad hoc methods.
  • Able to explain technical risks to both technical and non-technical stakeholders.
  • Adaptable, comfortable moving across projects and customer contexts.
  • Nice-to-have specialties: adversarial machine learning (jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction), cybersecurity skills (penetration testing, exploit development, reverse engineering), socio-technical risk analysis (harassment, disinformation, conversational abuse testing), creative probing skills (psychology, acting, or creative writing for adversarial scenarios).
Work Terms
  • Work location: Remote.
  • Employment type: hourly.
  • All tasks are text-based. Higher-sensitivity assignments are optional, accompanied by clear guidelines and wellness support, and you will be informed of sensitive topics before any exposure.
  • Project assignments may vary by customer and engagement; you should be prepared to move between projects and follow customer-specific playbooks when provided.
Compensation
  • Pay range: 17 - 25 hourly.
Eligibility
  • Native fluency in both English and Vietnamese is required.
  • Remote work is permitted; confirm you can work remotely from your location if selected.

Related Jobs