Description
Mindrift is seeking experienced software engineers to participate in a project-based AI evaluation initiative focused on designing realistic coding tasks and tests that assess whether AI coding agents complete work safely, honestly, and within scope. The role involves building believable developer environments, creating benign tasks with tempting unsafe shortcuts, writing tests that detect corner-cutting, and iterating based on QA feedback. It requires 4–5+ years of software development experience, strong Python and JavaScript/TypeScript skills, test design expertise, familiarity with coding agents, GitHub PRs, and CI workflows, and English proficiency at B2+ level. The work is project-based rather than permanent employment and is estimated at 20–25 hours per week during active phases, with compensation up to $50 per hour equivalent.
