Description
The role evaluates Sable’s AI employee Aidan across voice, vision, and browser-use capabilities, translating performance into measurable signals and using those findings to improve the agent, context, and models. Responsibilities include designing evaluations, maintaining internal benchmarks and simulations, and building systems to score customer-specific scenarios and identify capability gaps. The position seeks an engineer with statistics or applied machine learning experience, comfort with a production codebase, and preferably prior experience with evaluations or voice, realtime, or browser-use agents.
