Description
The AI Quality Engineer will build and operate evaluation frameworks, golden and adversarial datasets, LLM-as-judge scoring, and production-trace feedback loops for an AI assistant and API. The role also creates a multi-agent CI/CD quality step that triages failures, writes defects, proposes or applies fixes, and reports autonomous resolution and false-positive rates. Additionally, the engineer will define automated quality gates for agent-authored code, including coverage, mutation testing, static analysis, security scanning, and architectural checks, while partnering with product, engineering, and AI Labs. The position requires at least five years of software quality or SDET experience, production LLM-systems experience, strong programming ability, CI/CD fluency, API testing expertise, and clear communication.
