Description
The role builds infrastructure, tooling, and pipelines for evaluating Siri reliably at scale. It requires strong programming in a compiled language, Python scripting, computer science fundamentals, adaptability, and cross-team communication, with an M.S. or B.S. in a relevant field or equivalent experience. Preferred qualifications include experience with test or evaluation environments, ML or agent-based system evaluation, large-scale test infrastructure, and production Swift and Python.
