Enter a job title or keyword

Sr. Machine Learning Engineer, Foundation Models Inference Cloud OS & Inference

Apple


Job Location:

Paris, TX - USA

Monthly Salary: Not provided by the employer
Posted: 29 September 2026 (Yesterday)
Application Deadline: 27 December 2026
Vacancies: 1 Vacancy

Job Summary

We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant secure and private inference stack that powers Siri AI Apple Intelligence and Apps that are powered with the largest foundation models. Our systems serve billions of queries daily across Siri AI Apple Intelligence Apple Search Apple Music Apple TV App Store iMessage Photos Camera Spotlight u0026 Safari at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimise language vision and speech models with billions of parameters using state-of-the-art techniques and ship them at Apple scale. This is a rare opportunity to directly shape how AI reaches billions of people worldwide.

You will work at the intersection of research and production partnering closely with the Foundation Model Research team and our external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment. You will own hard problems in inference efficiency hardware/software codesign systems architecture and tooling and help set the technical direction for the engineers around you. This role sits within CloudOS and Private Cloud Compute (PCC) Apples purpose-built privacy-preserving cloud infrastructure for AI workloads. PCC represents a first-of-its-kind approach to running foundation models in the cloud with verifiable privacy guarantees and CloudOS is the systems foundation that makes it possible. You will be building and optimising inference systems on top of this infrastructure working closely with platform and security teams to deliver both performance and trust at scale.

Partner with the Foundation Model Research team and our external partners to optimise inference for the latest model architectures across language vision and and ship production-grade inference systems serving millions of customers in real profiling tools and simulators to identify and resolve performance bottlenecks across different hardware configurations and use technical decisions on high-throughput low-latency serving at supercomputing and grow engineers across the organisation.

Experience leading complex ambiguous Machine learning projects end to end. nProficiency in PyTorch or JAX nExperience working with Inference frameworksnExperienced in Python / Rust / Go lang or similar programming languagesnProficiency in deploying applications on cloud platforms (AWS GCP or equivalent) using K8S and docker.

Hands-on experience with LLM inference knowledge of GPU or TPU programming building and operating high-throughput services at large distributed building productions systems in Go or knowledge of deep learning architectures including Transformers encoder/decoder models and multimodal with inference optimization frameworks such as TensorRT-LLM vLLM SGLang TGI or Nvidia Triton in Computer Science Machine Learning Artificial Intelligence Data Science or a related field

Required Experience:

Senior IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile