Skip to content
NewEligibility scores are liveSee if you qualify
Science & researchOpen · checked today

AI Safety Experts — English & Assamese

We are looking for AI Safety Experts — English & Assamese candidates for a project delivered through Mercor.

What you'll do

  • Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing.
  • Document reproducibly by producing reports, datasets, and attack cases that customers can act on.

What you need

  • Bring strong judgment about language and content, with the ability to explain why an AI response is accurate, complete, and appropriate.
  • Be rigorous and notice subtle errors, inconsistencies, and gaps.
  • Be structured and work to guidelines and quality standards consistently.
  • Be communicative, explaining reasoning clearly to technical and non-technical audiences.
  • Be adaptable and thrive moving across projects, task types, and customers.

Nice to have

  • Experience with adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, and model extraction.
  • Background in cybersecurity, including penetration testing, exploit development, and reverse engineering.
  • Familiarity with socio-technical risk, including harassment/disinformation probing, abuse analysis, and conversational AI testing.
  • Skills in creative probing, psychology, acting, or writing for unconventional adversarial thinking.

Who you work with

Project and contracting process: Mercor. Applications continue on the provider's website.

More science & research roles

See all

$16–$22/hr

Mercor · Worldwide

Apply