Description
Mindrift is seeking a specialist to join their team to build a dataset for evaluating AI coding agents. The role involves creating challenging tasks and defining evaluation criteria within realistic simulated environments. This includes designing prompts, crafting test cases, and iterating on tasks based on feedback. The focus is on understanding where models fail and refining evaluations to accurately assess AI performance.
