Site Reliability Engineer, Apple Data Platform Big Data Platform

Apple


Job Location:

Austin, TX - USA

Monthly Salary: Not Disclosed
Posted on: 12 hours ago
Vacancies: 1 Vacancy

Job Summary

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apples long-held passion for combining art and technology. These are the people who power the App Store Apple TV Apple Music Apple Podcasts and Apple Books at extensive scale meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 ASE the Apple Data Platform SRE team keeps a massive multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure automation and customer success running incident response providing hands-on support to internal teams and partnering with developers to make cutting-edge services like Spark Flink Airflow Trino Notebooks and LLM-based agent platforms reliable at scale.n

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple while specialising in the big data engines and catalog/governance layers that power analytics and data engineering across the company. As an SRE on Apple Data Platform youll operate and support the teams full portfolio from ML/AI platform services to multi-cloud infrastructure and grow into the teams go-to expert for big data platform services including Spark Flink Airflow Trino Notebooks REST Catalog services (such as Glue Catalog) and data governance. Just as importantly youll be a first point of contact for the internal customers who rely on these services daily someone who can translate a confusing error or a vague support request into a clear diagnosis and a fast looking for a self-motivated engineer who thrives on ownership someone who wants a set of services to call their own the autonomy to drive their reliability roadmap and the collaborative instinct to keep that work aligned with the teams broader direction. If you love solving hard operational problems take genuine satisfaction in helping frustrated customers get unblocked and want a front-row seat to how Apples data engineering platform scales this role offers real room to grow your scope and impact over time.n

Operate monitor and triage production and non-production environments across the ADP portfolio data processing ML/AI and multi-cloud in a rotating on-call schedule across supported services including occasional weekday and weekend the operational health of big data platform services as SME driving reliability support and customer guidance for Spark Flink Airflow Trino Notebooks REST Catalog and governance as a primary point of contact for internal customers via Slack clearly communicating status root cause and next steps during active triage and resolve customer-reported service issues and support tickets prioritizing based on customer impact and with dev teams across time zones to onboard new services understanding architecture then designing monitoring alerting and dashboards (Prometheus Grafana Splunk).nBuild automation and self-healing tooling that reduces manual toil and scales the teams operational escalate and resolve production issues to protect platform reliability and customer customer success by helping internal teams understand platform capabilities and adopt tools with SRE and dev partner teams engineering and program management to align execution with team and org goals.

Bachelors Degree in Computer Science an engineering-related field or equivalent related experience.n1-4 years in a Site Reliability Engineering DevOps or Infrastructure-focused in Python; working knowledge of Golang a understanding of one or more Big Data technologies (Spark Flink Airflow Trino Notebooks).nExperience with Kubernetes and at least one major cloud provider (AWS or GCP).nExcellent written and verbal communication skills with the ability to explain technical issues clearly to non-expert grounding in SRE principles with prior on-call production-support or customer-facing support role experience.

Experience with REST Catalog services (e.g. Glue Catalog) and data governance experience in a customer-facing or technical support role with a demonstrated passion for customer with observability tooling: Prometheus Grafana Splunk knowledge of CI/CD pipelines and deployment with S3 and cloud storage/networking with data pipeline orchestration and workflow scheduling track record of automating manual operations through scripting or curiosity and a drive to keep learning for yourself your team and the org.n

Required Experience:

IC

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apples long-held passion for combining art and technology. These are the people who power the App Store Apple TV Apple Music Apple Podcasts and Apple Books at extensive scale meeting high expectations to deliver a hug...

About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile