Site Reliability Engineer, Apple Data Platform Multi-Cloud Infrastructure
Austin, TX - USA
Job Summary
This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple while specialising in the cloud infrastructure that underpins all of it. As an SRE on Apple Data Platform youll operate and support the teams full portfolio from big data pipelines to ML/AI platform services and grow into the teams go-to expert for multi-cloud infrastructure including AWS core services (IAM EKS RDS S3 VPC networking autoscaling EBS) Kubernetes administration at scale and the Infrastructure-as-Code and GitOps tooling that keeps it all reconciled and reliable. Youll be the person other engineers turn to when an IAM policy misfires a cluster hits a scheduling wall or a GitOps reconciliation drifts out of sync and the driving force behind making those failure modes rarer over looking for a self-motivated engineer who thrives on ownership someone who wants a set of services to call their own the autonomy to drive their reliability roadmap and the collaborative instinct to keep that work aligned with the teams broader direction. If you love going deep on cloud and Kubernetes internals enjoy being the trusted expert customers turn to and want a front-row seat to Apple Data Platforms multi-cloud evolution this role offers real room to grow your scope and impact over time.n
Operate monitor and triage production and non-production environments across the ADP portfolio data processing ML/AI and multi-cloud in a rotating on-call schedule across supported services including occasional weekday and weekend the operational health of multi-cloud infrastructure as SME driving reliability support and customer guidance for AWS services EKS clusters and cross-cloud Slack-based support to internal customers; screen triage and resolve service related production incidents involving IAM permission errors storage quota limits control-plane/data-plane namespace separation and cluster-wide with dev teams across time zones to onboard new services understanding architecture then designing monitoring alerting and dashboards (Prometheus Grafana Splunk).nMaintain and evolve Infrastructure-as-Code (Crossplane Terraform) and GitOps (Flux) workflows troubleshooting state drift and reconciliation automation and self-healing tooling that reduces manual toil and scales the teams operational escalate and resolve production issues to protect platform reliability and customer with SRE and dev partner teams engineering and program management to align execution with team and org goals.n
Minimum Qualificationsnn* Bachelors Degree in Computer Science an engineering-related field or equivalent related experience.n1-4 years in a Site Reliability Engineering DevOps or Infrastructure-focused in Python; working knowledge of Golang a hands-on AWS experience: IAM (roles policies permission boundaries KMS) EKS RDS S3 VPC networking/endpoints autoscaling groups administration experience RBAC node/pod scheduling autoscalers PriorityClasses/PDBs and troubleshooting cluster-wide communication skills and composure under pressure during grounding in SRE principles with prior on-call or production-support experience.n
Experience with Infrastructure-as-Code (Crossplane and/or Terraform) including debugging state drift and composition/controller with GitOps workflows (Flux or similar) HelmRepository/reconciliation troubleshooting and Helm chart -cloud exposure (GCP) parity and migration scenarios are emerging areas of with Splunk for log pipeline debugging (e.g. fluent-bit).nFamiliarity with Spark/Flink running on Kubernetes (executor scheduling node affinity).nComfort with GitHub PR review workflows in an infrastructure-as-code / GitOps track record of automating manual operations through scripting or curiosity and a drive to keep learning for yourself your team and the org.n
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more