Principal Data Engineering Lead Services Special Project

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not Disclosed
Posted on: 10 hours ago
Vacancies: 1 Vacancy

Job Summary

At Apple great ideas have a way of becoming phenomenal products services and customer experiences very quickly. Our team is building a massive real-time platform that transforms continuous streams of multimodal data (including structured image and log data) into an intelligent searchable foundation. nnWe are seeking a Principal Data Engineer to lead and drive not only our teams data processing systems but also to partner at a larger scale coordinating and synching strategically with other business groups and organizations within Apple. nn

We are seeking a Principal Data Engineering Lead with deep expertise in ETL/ELT data architecture and applied ML pipelines to drive the design build and operations of this infrastructure. As a key member of our team you will be responsible for driving critical decisions and operations across the entire system while aligning strategically across Apple.

Build and implement batch and streaming ETL/ELT pipelines that ingest process and model data from diverse sources including unstructured media and real-time event streams ensuring high reliability performance and and maintain Kafka-based ingestion and processing pipelines ensuring reliable data delivery across services and into the data robust logical and physical data models with a focus on dimensional modeling versioning and storage patterns (e.g. Parquet ORC) optimized for ingest reporting and operational use and enforce data quality checks SLAs and observability standards to ensure data is accurate timely versioned and trusted by and enrich raw signals with metadata and attribution to power downstream use cases such as analytics billing planning and standard methodologies for data lineage metadata management schema governance versioning and security in alignment with Apples standards for data protection and solutions that include logging anomaly detection data validation cleaning and transformation with strong emphasis on monitoring debuggability and continuous closely with ML engineers data scientists platform teams and leadership to translate requirements into scalable reliable data advance the teams data stack including tooling frameworks and standards for development testing deployment and our team with other Apple teams strategically participating in larger scale discussions and deliverables across our ecosystem. n

Masters Degree n12 years of experience in data engineering including building and maintaining large-scale ETL/ELT data pipelinesnProficiency in data modeling especially dimensional modeling and designing schemas optimized for analytics and reportingnExperience with leveraging databases including SQL/NoSQL Databases (including Postgres / Cassandra / Redis)nStrong experience with distributed data processing frameworks including Apache Spark nStrong experience with Parallel processing frameworks: BigTable/HadoopnStrong software engineering fundamentals and proven experience with Scala JavanHands-on experience with Apache Kafka Iceberg and Flink. nExperience with workflow orchestration tools including Apache Airflow and BeamnExperience with AWS: e.g. S3 EMR Lambda Glue Redshift BigQuery Kinesis or similar servicesnExperience with Analytics frameworks including Trino (Presto BigQuery Snowflake)nHands-on experience with big data lake architecturesnExperience with containerization and orchestration (Docker Kubernetes/EKS) and CI/CD tooling including Jenkins nExperience in Python and PySparknFamiliarity with graph databases such as TigerGraphnExperience building pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformationnHands-on experience deploying serving and optimizing LLMs or ML models directly in the production inference runtimes/compilers (ONNX Runtime TensorRT/TensorRT-LLM) and serving frameworks (Triton vLLM TorchServe or similar). nExperience tuning batching KV-cache and GPU utilization for low-latency high-throughput real-time inference in a data pipelinenKnowledge of data governance principles data security best practices and data privacy regulationsnProven experience delivering a consumer-oriented solution by participating at every stage of the development life-cycle. nExcellent communication skills and a collaborative mindset with past experience presenting and partnering with VP and C level decision makers. n

Experience with data versioning tools and frameworks (e.g. DVC Delta Lake)nExperience storing/serving embeddings (e.g. pgvector Milvus FAISS)

Required Experience:

Staff IC

At Apple great ideas have a way of becoming phenomenal products services and customer experiences very quickly. Our team is building a massive real-time platform that transforms continuous streams of multimodal data (including structured image and log data) into an intelligent searchable foundation....

About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile