Description
Vals AI is hiring an LLM Benchmarking Engineer to evaluate new large language model releases across its benchmark suite, analyze model error modes, maintain model integrations, and support benchmark infrastructure. The role requires Python expertise, strong engineering fundamentals, collaboration in development sprints, and the ability to work intensively during model releases. It is an in-person position based in San Francisco, with relocation or transportation support, health and dental insurance, a 401K plan, unlimited PTO, and a housing stipend.
