Description
Office Hours is hiring a Software Engineer, Benchmarking to build and maintain the systems that run AI model evaluations. The role prepares benchmark datasets, creates evaluation pipelines and containerized environments, supports model experiments, develops scoreboards and leaderboards, and builds analysis and data-review tools. It requires at least four years of professional experience, strong Python, data rigor, and familiarity with Docker; the position is available in San Francisco, New York City, or fully remotely, with a base salary of $160,000-$210,000 plus equity and benefits.
