Apple Services Engineering (ASE) Compute Software Engineering Manager
Cupertino, CA - USA
Job Summary
Apple Service Engineering (ASE)s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apples data centers. You will manage a team that operates core compute controllers proxy services job execution agents and supporting infrastructure across multiple geographies ensuring platform availability reliability and performance at Apple will drive strategic initiatives spanning multi-datacenter capacity planning incident management release engineering observability and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation establishing SLOs driving production readiness and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage improve operational workflows drive capacity efficiency and enhance team productivity across all domains.n
Champion AI-powered tooling and automation to improve incident triage reduce operational toil drive capacity efficiency and accelerate engineering workflowsnLead mentor and grow a team of Software and SRE engineers across multiple geographies fostering a culture of ownership collaboration and continuous improvementnEstablish and maintain a sustainable 24/7 on-call rotation with clear escalation paths severity definitions and response time SLAs across US and UK locationsnOversee release engineering and deployment automation including CI/CD pipelines canary deployments and zero-downtime rollout strategiesnManage infrastructure modernization initiatives including Kubernetes control plane operations database migrations and configuration management evolutionnDrive incident management excellence including post-incident reviews preventive measures and production readiness reviews for all releases
5 years of experience managing infrastructure SRE or platform engineering teams operating large-scale distributed systemsnProven track record of building and leading on-call organizations with structured incident management escalation procedures and post-incident review processesnStrong technical background in cloud infrastructure compute orchestration and bare metal provisioning at scalenExperience with Kubernetes OpenStack KVM/hypervisor technologies and Infrastructure as Code tools (Chef Ansible Terraform or Salt)nDeep understanding of SRE principles including SLOs error budgets capacity planning and release engineeringnExcellent verbal and written communication skills with the ability to influence across teams and levelsnDemonstrated ability to recruit develop and retain high-performing engineering talent
Hands-on experience leveraging AI and machine learning to improve operational efficiency incident management or infrastructure automationnExperience managing or scaling batch compute job scheduling or HPC platformsnProficiency in Go or Python with a strong automation-first mindsetnFamiliarity with observability stacks (Prometheus Grafana distributed tracing) and centralized logging at scalenExperience operating large-scale multi-tenant Infrastructure as a Managed ServicenExperience managing geographically distributed teams and follow-the-sun on-call modelsnTrack record of driving capacity efficiency initiatives resulting in measurable cost optimization
Required Experience:
Manager
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more