We are looking for Senior Software Engineer - AI Evaluation / Coding Agents candidates for a project delivered through the hiring partner.
What you'll do
- Evaluate AI-generated code and solutions across real-world software repositories
- Review agent behavior, tool usage, and code changes for correctness and quality
- Identify technical errors, weak approaches, and recurring model failure modes
- Compare model outputs and explain why one solution is better than another
- Create and refine rubrics and evaluation criteria for coding tasks
- Produce high-quality evaluation and preference data used to improve coding models
- Build and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows
- Synthesize findings from data work into clear write-ups, updates, and recommendations for the team
- Collaborate closely with researchers and engineers to translate qualitative judgment into scalable processes
- Share clear, actionable findings with AI researchers and engineers
What you need
- 5+ years of hands-on software engineering experience
- Strong proficiency in Python, TypeScript/JavaScript, Go, or another major production language
- Experience working in substantial real-world codebases
- Strong code-review skills and technical judgment
- Ability to clearly explain why an implementation is correct, incorrect, or could be improved
- Strong written communication
- Experience using modern LLMs or AI coding tools
Nice to have
- Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training
Expertise
Who you work with
Project and contracting process: the hiring partner. Applications continue on the provider's website.

