Description
Mindrift is seeking a specialist to join a project-based team focused on evaluating AI coding agents. The role involves creating challenging tasks and defining evaluation criteria within simulated environments to test how well AI models handle real-world developer tasks. Responsibilities include designing prompts, crafting tasks, writing tests, and iterating based on feedback. This is not data labeling, prompt engineering, or direct coding.
