Skip to main content

Evaluations Engineer at Vals Smith

Department: Engineering & Research

Compensation

$140,000 – $185,000/yr

Setup
On-site
Location
San Francisco, California
Type
Full-time
Level
not_specified
Posted

Description

Vals AI is hiring an LLM Benchmarking Engineer to evaluate new large language model releases across its benchmark suite, analyze model error modes, maintain model integrations, and support benchmark infrastructure. The role requires Python expertise, strong engineering fundamentals, collaboration in development sprints, and the ability to work intensively during model releases. It is an in-person position based in San Francisco, with relocation or transportation support, health and dental insurance, a 401K plan, unlimited PTO, and a housing stipend.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation