Description
This project-based opportunity focuses on creating challenging tasks and evaluation criteria for AI coding agents. The role involves designing prompts, defining success metrics, and developing tests to assess how AI models handle real-world developer tasks within simulated environments. It's not data labeling, prompt engineering, or direct code writing.
