Description
The role owns technical testing of agent outputs, including accuracy, turnaround-time impact, and regression testing as agents are iterated. Responsibilities include building test suites against golden-record datasets, tracking accuracy and completion-time metrics, running regression tests when agent or prompt configurations change, and providing technical evidence to the human QA function for final sign-off on regulated activity. The role requires at least five years of QA or test engineering experience, ideally including ML/LLM output validation, and experience defining test datasets and accuracy benchmarks for document or text-processing systems.
