Description
METR is hiring an AI Evaluation Engineer to run frontier-model evaluations, integrate models into agent scaffolds, debug and improve evaluation infrastructure, manage concurrent evaluation projects, and communicate results to researchers, regulators, and other stakeholders. The role requires software engineering, infrastructure, debugging, performance optimization, attention to detail, and strong communication; research understanding, external communication, project management, and writing are preferred. The position is hybrid in Berkeley, California, with technical team members working in the office three to five days per week, and offers base salary of $328,000–$578,000 depending on seniority.
