Description
Mindrift is seeking a specialist to join their team to build a dataset for evaluating AI coding agents. The role involves creating challenging tasks and evaluation criteria within realistic simulated environments, designing tasks from intermediate states of these environments, writing tests to verify agent solutions, and iterating on tasks and tests based on QA feedback. This is a project-based opportunity focusing on testing, evaluating, and improving AI systems.
