Enter a job title or keyword

Sr. Machine Learning Engineer, Speech LLM Evaluation

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not provided by the employer
Posted: 10 September 2026 (2 days ago)
Application Deadline: 8 December 2026
Vacancies: 1 Vacancy

Job Summary

Join the team redefining what a deeply personal and integrated assistant can part of the Siri organization you will help shape one of the worlds most widely used AI assistants powered by our next-generation of Apple Intelligence with capabilities like personal context understanding and on-screen awareness built with privacy from the ground up. Your work will have direct meaningful impact for users across iOS iPadOS macOS watchOS and Speech Evaluation team sits at the center of Apples ASR TTS and real-time conversational AI efforts partnering directly with the modeling teams. Were growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether theyre ready. Youll help define how Apple measures a new class of models that listen speak and is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design shipping technology that is centered around users and their needs.

This role owns the data and metrics foundation for evaluating speech LLMs (e.g. real-time speech understanding and generation models) across accuracy robustness and conversational quality. Youll build and curate evaluation datasets that reflect real usage from personalized named-entity queries to multi-turn fluid conversations and design the metrics and automated judges that turn model outputs into actionable trustworthy signal. Youll work closely with modeling infrastructure and product partners to make sure every new model is evaluated quickly consistently and at the right level of rigor before it reaches customers.n

Designs and curates audio evaluation datasets that represent real-world usage including personalized multilingual and conversational and implements evaluation metrics for audio LLMs spanning accuracy robustness and conversational/generation automated evaluation pipelines and LLM-as-judge tooling to scale audio model assessment without sacrificing model evaluation results to identify accuracy gaps regressions and opportunities for hillclimbing and communicates findings to modeling with human-evaluation programs to design rating protocols and validate that automated metrics correlate with human with infrastructure teams to integrate new evaluation sets and metrics into shared evaluation methodology for new audio LLM capabilities as they emerge adapting existing frameworks to novel model behaviors.

Bachelors degree in Computer Science Electrical Engineering or a related field or equivalent practical building or working with text speech or audio evaluation pipelines and in Python and experience building data processing pipelines at curating or annotating datasets for machine learning evaluation or knowledge of statistics as applied to measuring model performance and interpreting evaluation with large language model evaluation techniques including automated (LLM-as-judge) and human evaluation written and verbal communication skills with the ability to explain evaluation results to both technical and non-technical audiences.

Experience evaluating audio-native or multimodal (speech-in speech-out) large language designing or running human evaluation studies (e.g. side-by-side comparisons MOS ratings) at with personalization and named-entity evaluation challenges in speech with multilingual or international audio dataset with distributed data processing frameworks (e.g. Spark) for large-scale audio dataset record or demonstrated contributions in speech audio ML or NLP evaluation.

Required Experience:

Senior IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile