Description
This project-based opportunity focuses on creating challenging tasks and evaluation criteria for AI coding agents. The goal is to build a dataset to test, evaluate, and improve AI systems by simulating realistic developer environments and designing tasks that challenge frontier models. Participants will iterate on tasks and tests based on feedback to ensure fair and robust evaluations.

