Description
This project-based opportunity focuses on creating challenging tasks and evaluation criteria to assess AI coding agents. The role involves designing prompts, defining success metrics, and developing tests to validate agent solutions. It's crucial to iteratively refine tasks based on feedback to ensure fairness and robustness. The project does not involve data labeling, prompt engineering, or direct code writing.

