At Apple we dont just build products we create transformative experiences that have reshaped entire industries. Our innovation is driven by the diversity of our people and their ideas inspiring everything we do. Imagine the impact you could make. Join Apple and help us leave the world better than we found it. The ML Infrastructure team is responsible for managing Apples largest ML compute platform multi-cloud storage abstraction and caching platform which supports critical machine learning training workloads that power user-facing features across the Apple ecosystem. Operating across both first-party and third-party cloud environments brings complex and unique challenges. As a Site Reliability Engineer (SRE) on the ML Infrastructure team youll be expected to address these challenges through a strong foundation in cloud object storage data analysis automation collaboration and advanced expertise in Kubernetes. Our team oversees the full infrastructure stack from low-level nodes to the complete network architecture ensuring our platform remains highly available resilient and efficient at scale.
We are seeking an experienced Software and Systems Engineer to join our dynamic team. This role demands a proactive mindset technical excellence and a collaborative spirit. The ideal candidate will demonstrate: Strong critical thinking and a high degree of individual accountability Effective communication and collaboration skills A genuine passion for Infrastructure as a Service (IaaS) A commitment to automation and operational efficiency Ownership of projects from design through delivery A solutions-oriented approach coupled with the ability to gain alignment on technical direction Consistent and timely execution of design implementations aligned with project objectives The ability to provide constructive technical feedback fostering team-wide growth and continuous improvementn
Participates in a rotating on-call schedule including occasional weekendncoverage when necessarynCurrently headquartered in Cupertino with active expansion in Bangalore tonsupport global operations across time zonesnLeverages a diverse stack including open-source tools commercial solutionsnand internally developed systemsnEncourages open dialogue values strong ideas and recognizes impactful resultsn
5 years experience in building operating and scaling a large application in anprivate public or hybrid cloud environmentnDeep expertise in Kubernetes with hands-on experience using platforms suchnas Google Kubernetes Engine (GKE) or Amazon Elastic Kubernetes Servicen(EKS)nProficient in designing developing and releasing code in languages such asnPython Go or RustnPractical experience with object storage technologies including Amazon S3nor Google Cloud Storage (GCS)nStrong background in designing and troubleshooting complex networkingnissues in both public and private cloud infrastructuresnSolid understanding of Linux internals standard networking protocols andndistributed systems architecturen
Proven drive to automate manual operations and enhance processes throughncontinuous iterationnStrong understanding of best practices for deploying large-scale distributednapplicationsnHands-on experience managing diverse system environments usingnconfiguration management tools or software delivery platforms such asnSpinnaker Helm or FluxnDemonstrated expertise in deploying supporting and monitoring both newnand existing services platforms and application stacksnSolid familiarity with container orchestration and management usingnKubernetesn
Required Experience:
Senior IC
At Apple we dont just build products we create transformative experiences that have reshaped entire industries. Our innovation is driven by the diversity of our people and their ideas inspiring everything we do. Imagine the impact you could make. Join Apple and help us leave the world better than w...
At Apple we dont just build products we create transformative experiences that have reshaped entire industries. Our innovation is driven by the diversity of our people and their ideas inspiring everything we do. Imagine the impact you could make. Join Apple and help us leave the world better than we found it. The ML Infrastructure team is responsible for managing Apples largest ML compute platform multi-cloud storage abstraction and caching platform which supports critical machine learning training workloads that power user-facing features across the Apple ecosystem. Operating across both first-party and third-party cloud environments brings complex and unique challenges. As a Site Reliability Engineer (SRE) on the ML Infrastructure team youll be expected to address these challenges through a strong foundation in cloud object storage data analysis automation collaboration and advanced expertise in Kubernetes. Our team oversees the full infrastructure stack from low-level nodes to the complete network architecture ensuring our platform remains highly available resilient and efficient at scale.
We are seeking an experienced Software and Systems Engineer to join our dynamic team. This role demands a proactive mindset technical excellence and a collaborative spirit. The ideal candidate will demonstrate: Strong critical thinking and a high degree of individual accountability Effective communication and collaboration skills A genuine passion for Infrastructure as a Service (IaaS) A commitment to automation and operational efficiency Ownership of projects from design through delivery A solutions-oriented approach coupled with the ability to gain alignment on technical direction Consistent and timely execution of design implementations aligned with project objectives The ability to provide constructive technical feedback fostering team-wide growth and continuous improvementn
Participates in a rotating on-call schedule including occasional weekendncoverage when necessarynCurrently headquartered in Cupertino with active expansion in Bangalore tonsupport global operations across time zonesnLeverages a diverse stack including open-source tools commercial solutionsnand internally developed systemsnEncourages open dialogue values strong ideas and recognizes impactful resultsn
5 years experience in building operating and scaling a large application in anprivate public or hybrid cloud environmentnDeep expertise in Kubernetes with hands-on experience using platforms suchnas Google Kubernetes Engine (GKE) or Amazon Elastic Kubernetes Servicen(EKS)nProficient in designing developing and releasing code in languages such asnPython Go or RustnPractical experience with object storage technologies including Amazon S3nor Google Cloud Storage (GCS)nStrong background in designing and troubleshooting complex networkingnissues in both public and private cloud infrastructuresnSolid understanding of Linux internals standard networking protocols andndistributed systems architecturen
Proven drive to automate manual operations and enhance processes throughncontinuous iterationnStrong understanding of best practices for deploying large-scale distributednapplicationsnHands-on experience managing diverse system environments usingnconfiguration management tools or software delivery platforms such asnSpinnaker Helm or FluxnDemonstrated expertise in deploying supporting and monitoring both newnand existing services platforms and application stacksnSolid familiarity with container orchestration and management usingnKubernetesn
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar
... View more