Description
This project-based opportunity focuses on creating challenging tasks and evaluation criteria for AI coding agents. The goal is to build a dataset to assess how well AI models handle real-world developer tasks within realistic simulated environments. Responsibilities include designing tasks, crafting prompts, defining success metrics, and iterating on evaluations based on feedback. It's important to note that this isn't data labeling, prompt engineering, or direct code writing.

