We are looking for AI Safety Experts — English & Gujarati candidates for a project delivered through the hiring partner.
What you'll do
- Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
- Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
- Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
- Document reproducibly: produce reports, datasets, and attack cases customers can act on
What you need
- Strong judgment about language and content: ability to determine if an AI response is accurate, complete, and appropriate, and explain why
- Rigorous: notice subtle errors, inconsistencies, and gaps that others skim past
- Structured: work to guidelines and quality standards consistently, not ad hoc
- Communicative: explain reasoning clearly to technical and non-technical audiences
- Adaptable: thrive moving across projects, task types, and customers
Nice to have
- Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
- Cybersecurity: penetration testing, exploit development, reverse engineering
- Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
- Creative probing: psychology, acting, writing for unconventional adversarial thinking
Who you work with
Project and contracting process: the hiring partner. Applications continue on the provider's website.

