Description
Apple is hiring a Speech Evaluation Engineer to build the data and metrics foundation for evaluating audio large language models. The role involves curating real-world audio evaluation datasets, defining accuracy, robustness, and conversational-quality metrics, developing automated evaluation pipelines and LLM-as-judge tooling, analyzing model results, and partnering with modeling, infrastructure, and human-evaluation teams. The position requires a relevant bachelor's degree or equivalent experience, prior experience with speech or audio evaluation pipelines, Python and data-processing skills, dataset curation experience, statistical knowledge, and familiarity with LLM evaluation methods.
