We are looking for AI Safety Experts — English & Thai candidates for a project delivered through Mercor.
What you'll do
- Red team conversational AI models and agents, focusing on jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
- Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
- Apply structure to testing by following taxonomies, benchmarks, and playbooks for consistent testing.
- Document reproducibly by producing reports, datasets, and attack cases that customers can act on.
What you need
- Prior red teaming experience, including AI adversarial work, cybersecurity, or socio-technical probing.
- Demonstrated curiosity and adversarial mindset, with an instinct to push systems to breaking points.
- Structured approach to testing, utilizing frameworks or benchmarks rather than random methods.
- Strong communication skills to clearly explain risks to both technical and non-technical stakeholders.
- Adaptability to thrive when moving across different projects and customers.
Nice to have
- Specialization in Adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
- Background in Cybersecurity, with experience in penetration testing, exploit development, or reverse engineering.
- Expertise in Socio-technical risk, such as harassment/disinformation probing, abuse analysis, or conversational AI testing.
- Creative probing skills, potentially from a background in psychology, acting, or writing, for unconventional adversarial thinking.
Expertise
Who you work with
Project and contracting process: Mercor. Applications continue on the provider's website.

