We are looking for AI Safety Experts — English & Bengali candidates for a project delivered through Mercor.
What you'll do
- Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
- Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
- Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
- Document reproducibly: produce reports, datasets, and attack cases customers can act on
What you need
- You bring strong judgment about language and content: you can tell whether an AI response is accurate, complete, and appropriate, and explain why
- You’re rigorous: you notice subtle errors, inconsistencies, and gaps that others skim past
- You’re structured: you work to guidelines and quality standards consistently, not ad hoc
- You’re communicative: you explain your reasoning clearly to technical and non-technical audiences
- You’re adaptable: you thrive moving across projects, task types, and customers
Nice to have
- Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
- Cybersecurity: penetration testing, exploit development, reverse engineering
- Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
- Creative probing: psychology, acting, writing for unconventional adversarial thinking
Who you work with
Project and contracting process: Mercor. Applications continue on the provider's website.

