Description
Mindrift is seeking a specialist to join a project-based team focused on evaluating AI coding agents. The role involves creating challenging tasks and defining evaluation criteria within simulated environments to assess how well AI models handle real-world developer tasks. Responsibilities include designing prompts, crafting test cases, iterating on tasks based on feedback, and ensuring evaluations are fair and robust.

