We are looking for SWE-Bench Task Auditor candidates for a project delivered through Mercor.
What you'll do
- Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models.
- Assess repository-level tasks, reference patches, test harnesses, and grading integrity.
- Provide clear, rubric-based written feedback.
What you need
- 3+ years professional software engineering
- Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
- Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
- Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Nice to have
- Familiarity with SWE-Bench (Verified) or similar repository benchmarks
- Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)
- Prior code-review or task-grading experience
Expertise
Who you work with
Project and contracting process: Mercor. Applications continue on the provider's website.

