We are looking for AI Safety Experts — English & Indonesian candidates for a project delivered through the hiring partner.
What you'll do
- Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
- Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
- Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
- Document reproducibly: produce reports, datasets, and attack cases customers can act on
What you need
- You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing)
- You’re curious and adversarial: you instinctively push systems to breaking points
- You’re structured: you use frameworks or benchmarks, not just random hacks
- You’re communicative: you explain risks clearly to technical and non-technical stakeholders
- You’re adaptable: thrive on moving across projects and customers
- Native fluency in English
- Native fluency in Indonesian
Nice to have
- Specialty in Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
- Specialty in Cybersecurity: penetration testing, exploit development, reverse engineering
- Specialty in Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
- Specialty in Creative probing: psychology, acting, writing for unconventional adversarial thinking
Expertise
Who you work with
Project and contracting process: the hiring partner. Applications continue on the provider's website.

