Enter a job title or keyword

Observability SRE Manager, Apple Services Engineering

Apple


Job Location:

Seattle, OR - USA

Monthly Salary: Not provided by the employer
Posted: 21 August 2026 (18 hours ago)
Application Deadline: 18 November 2026
Vacancies: 1 Vacancy

Job Summary

People at Apple dont just build products they craft the kind of experiences that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! nnJoin Apple and help us leave the world better than we found it. nnThe Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apples services (such as iCloud iTunes Siri and Maps). We are the foundation on which Apples software developers build the products that our customers love. nnApples observability and monitoring platforms are the nervous system behind the reliability of Apples cloud services giving thousands of engineers the visibility they need to detect diagnose and resolve issues before they impact customers. Our cloud monitoring platform analyzes billions of metrics per minute and is the first place incident responders turn when something goes wrong regardless of system scale or complexity. nnWere looking for a senior SRE leader to own and evolve this platform: the metrics logging tracing and alerting infrastructure that underpins operational excellence across Apple Services Engineering.

Youll set technical direction for reliability and operational excellence while mentoring engineers driving automation and partnering closely with software infrastructure finance and product teams to uphold the reliability of the platform whilst shipping improvements that matter at Apple is a senior leadership position that defines where our observability platform reliability scalability and performance goes next how our SRE practice evolves including how AI reshapes it and how we build the team and partnerships to get there. You will lead engineers solving reliability and scale problems few organizations encounter integrating monitoring seamlessly across disparate infrastructures hardware software application and network layers at a scale built to reach every user on the planet. Youll build a team culture that makes SRE sustainable rewarding and central to how Apple ships successful candidate has a strong aptitude for both technical leadership and people management with the ability to context-switch between strategic planning and tactical execution. You should be comfortable building and scaling teams driving complex cross-functional initiatives and thriving under pressure while creating an inclusive high-trust team culture where engineers do their best believe AI will fundamentally reshape how SRE is practiced from anomaly detection and root-cause analysis to capacity planning and toil elimination and were looking for a leader who shares that conviction and can drive that transformation across the organization.

Technical u0026 Operational LeadershipnnOwn the reliability availability and performance of Apples observability platform (metrics logging tracing alerting) and the self-service capabilities built on top of itnLead staging and production environments for the observability platform with the goal of maximizing availabilitynEnsure the platform accurately monitors the health of every application and piece of infrastructure across the Apple ecosystem the central nervous system that other engineering teams and incident responders reach for firstnDefine and drive the strategic roadmap for observability infrastructure in partnership with SRE engineering and product stakeholdersnEstablish and refine SRE practices including SLOs error budgets capacity planning scale testing disaster recovery and change managementnGuide deep dives into systemic and latent reliability issues spanning the full stack (hardware software application and network) partnering with software and systems engineers to drive fixes to resolutionnDrive incident response post-incident reviews and systemic improvements that reduce operational toilnChampion automation to eliminate manual processes through tooling self-service platforms and APIs for internal customersnManage on-call rotations and ensure sustainable well-supported operational coveragenDrive standardization of monitoring and troubleshooting methodology across embedded SREs and services throughout the organizationnRepresent the SRE organization in design reviews and operational readiness exercises for new and existing servicesnPartner with software engineering and architecture teams to influence system design for reliability scalability and operabilitynBalance technical debt reduction with feature development to maintain platform healthnnAI u0026 ModernizationnnDefine and execute a clear AI strategy for the SRE organization identifying high-impact opportunities where AI/ML tooling can reduce toil accelerate root-cause analysis and improve reliability outcomesnDrive adoption of AI-assisted tooling (copilots intelligent runbooks LLM-based diagnostics anomaly detection) into day-to-day SRE workflowsnBuild a culture where engineers actively experiment with AI tools and modern approaches to solve operational problemsnnPeople Leadership u0026 Org BuildingnnLead grow and mentor a team of Site Reliability Engineers conducting regular 1:1s performance reviews and career development discussionsnHire and build out the SRE team with a desire to develop engineers to meet both their career goals and the organizations goalsnBuild and scale a high-performing team through coaching clear expectations and psychological safetynFoster an inclusive team culture and mentor diverse talentnChampion engineering best practices for code quality system design and operational excellence across the broader organizationnnCross-Functional u0026 Executive CommunicationnnBuild strong partnerships across Apple negotiating priorities and aligning on shared goals with an Apple-first mindsetnRepresent the SRE perspective in cross-functional planning translating reliability requirements into architecture and investment decisionsnCommunicate effectively at the executive level presenting strategy trade-offs and progress to senior leadershipnDrive standardization and best practices across the observability and infrastructure landscapennn

5 years of engineering management experience leading SRE infrastructure or observability/monitoring teamsnExperience hiring and leading engineers with a desire to build grow and mentor a teamnDeep understanding of observability systems and practices: metrics logging tracing alerting SLOs error budgets and fault analysis at scalenStrong systems background comfortable troubleshooting across the full stack (network OS container runtime application)nExperience operating large-scale multi-tenant distributed systems in production including Kubernetes environmentsnPractical solid knowledge of shell/bash scripting and at least one higher-level production language (Python preferred; Go Java or Scala also valued)nDemonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operationsnTrack record of building high-performing teams through coaching clear expectations and psychological safetynDemonstrated ability to drive cross-functional initiatives to completion and communicate at the executive levelnBachelors or Masters degree in Computer Science Engineering or related field or equivalent experience

Deep familiarity with the Prometheus ecosystem and cloud-native observability stacks (Thanos Splunk OpenTelemetry or similar)nExperience with third-party cloud platforms (AWS GCP or Azure) and infrastructure as code (Terraform Ansible)nComfortable with open-source configuration management and orchestration tools (Helm Puppet Spinnaker)nDemonstrable knowledge of TCP/IP HTTP web application security and multi-tier web application architecturesnExperience running infrastructure as an internal managed service with defined SLAsnFamiliarity with microservices architecture and container orchestration with Kubernetes at scalenBackground in capacity planning performance engineering or infrastructure architecturenTrack record of driving cultural and process transformation within SRE organizationsnExperience building or deploying AI-powered operational tooling (AIOps intelligent alerting automated diagnostics)nDeveloping and delivering multi-mode communications tailored to the unique needs of different audiencesnAnticipating and balancing the needs of multiple stakeholdersnMaking sense of complex high-quantity and sometimes contradictory information to solve problems effectivelynRebounding from setbacks and adversity when facing difficult situationsnKnowing the most effective and efficient processes to get things done with a focus on continuous improvementn

Required Experience:

Manager


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile