Description
The role owns the automated evaluation approach for AI-powered productivity and creative features, including multi-turn conversation and end-to-end agent workflow testing. Responsibilities include building adversarial and stress-test suites, developing evaluation frameworks and rubrics, aligning automated and human evaluation, integrating evaluation into development and release workflows, and communicating readiness assessments. The position requires a bachelor’s degree in a relevant field, at least four years of experience building or extending ML evaluation systems, experience with adversarial or red-teaming methodologies, and production or near-production Python and ML framework experience.
