We are looking for AI Safety Experts — English & Kannada candidates for a project delivered through Mercor.
What you'll do
- Red team conversational AI models and agents including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structure by following taxonomies, benchmarks, and playbooks to maintain consistent testing.
- Document findings reproducibly by producing reports, datasets, and attack cases that customers can act on.
What you need
- Possess strong judgment about language and content, with the ability to determine if an AI response is accurate, complete, and appropriate, and explain the reasoning.
- Demonstrate rigor by noticing subtle errors, inconsistencies, and gaps that others might overlook.
- Work in a structured manner, adhering to guidelines and quality standards consistently.
- Communicate clearly, explaining reasoning to both technical and non-technical audiences.
- Be adaptable and thrive when moving across different projects, task types, and customers.
Nice to have
- Experience with Adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, and model extraction.
- Background in Cybersecurity, including penetration testing, exploit development, and reverse engineering.
- Knowledge of Socio-technical risk, including harassment/disinformation probing, abuse analysis, and conversational AI testing.
- Skills in Creative probing, such as psychology, acting, or writing for unconventional adversarial thinking.
Who you work with
Project and contracting process: Mercor. Applications continue on the provider's website.

