Skip to content
NewEligibility scores are liveSee if you qualify
AI & dataOpen · checked today

AI Safety Experts — English & Malay

We are looking for AI Safety Experts — English & Malay candidates for a project delivered through Mercor.

What you'll do

  • Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to ensure testing consistency.
  • Document findings reproducibly by producing reports, datasets, and attack cases for customer action.

What you need

  • Prior red teaming experience, including AI adversarial work, cybersecurity, or socio-technical probing.
  • A curious and adversarial mindset, instinctively pushing systems to their breaking points.
  • A structured approach, utilizing frameworks or benchmarks for testing.
  • Strong communication skills to explain risks clearly to both technical and non-technical stakeholders.
  • Adaptability to thrive when moving across different projects and customers.

Nice to have

  • Specialty in Adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
  • Specialty in Cybersecurity, including penetration testing, exploit development, or reverse engineering.
  • Specialty in Socio-technical risk, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
  • Specialty in Creative probing, involving psychology, acting, or writing for unconventional adversarial thinking.

Expertise

Who you work with

Project and contracting process: Mercor. Applications continue on the provider's website.

More ai & data roles

See all

$17–$25/hr

Mercor · Worldwide

Apply