Description
The Machine Learning Engineer specializing in Evaluation will establish evaluation criteria, metrics frameworks, and quality standards for machine learning models powering Apple Wallet, Payments, and Commerce features. The role owns the full evaluation lifecycle, including adversarial test strategies, fairness and robustness testing, generative-model assessment, user-stratified benchmarks, and final model-quality sign-off. The engineer will collaborate with ML Engineering, Product, Privacy, and Legal teams to surface failure modes, guide model development priorities, and ensure reliable model launches for hundreds of millions of users.
