Principal Data Architect and Manager Service Special Projects

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not Disclosed
Posted on: 13 hours ago
Vacancies: 1 Vacancy

Job Summary

Were building the large-scale data foundation that powers private personalized experiences across Apple platforms. Our team designs and operates the systems that ingest unify and understand information at massive scale turning petabytes of data from many sources into a single high-quality richly structured representation. This foundation is what intelligent search and on-device experiences rely on and we build it with an uncompromising bar for data quality freshness and are looking for a Principal Data Architect and Manager to serve as both the senior technical authority and the people leader for our data platform.

As the Principal Data Architect and Manager on our team you will serve as both the senior technical authority and the people leader for our data platform. Youll define and own the end-to-end architecture of a real-time petabyte-scale data backbone: from ingestion through a multi-layered lakehouse to normalized serving layers that power downstream search ranking and on-device experiences. Youll also build grow and lead the team of data engineers who bring that architecture to is a hands-on principal role with multiple facets: you set the technical vision personally shape the hardest architectural decisions drive the roadmap through to production and manage mentor and grow the engineers executing against it. Your leverage comes equally from what you design and from the team you build.n

1. Architecture u0026 Design (Architect scope)nnDefine the end-to-end architecture of a multi-layered lakehouse on cloud object storage as the canonical layer partitioning strategy columnar formats (Parquet) open table formats (Iceberg or Delta) compaction and cost management at petabyte the data governance framework: schema registries lineage metadata management quality checkpoints and access controls spanning the full lifecycle from raw ingestion to normalized serving layers aligned with Apples privacy and security batch micro-batch and streaming ETL/ELT pipelines capable of handling structured semi-structured and unstructured multimodal data including image and other media with real-time metadata extraction schema augmentation and build and operate a fault-tolerant Apache Kafka streaming backbone including topic design schema evolution consumer-group topology and delivery-semantics guarantees across the architectural direction for entity resolution conflation and knowledge-graph construction at the scale of billions of frequently updated privacy as an architectural constraint not a compliance step: data minimization retention and deletion enforcement and data privacy constraints designed into the platform from the first 2. Technical Leadership u0026 Implementation (Lead scope)nnSet the technical roadmap for the data platform and drive the teams execution against it from architectural vision through to production the ingestion services that reliably land massive heterogeneous streams from various partners and own the data contracts with those the implementation of complex data transformations normalization augmentation enrichment with a strong bar for correctness consistency and analytical optimize pipeline performance reliability and cost evaluating trade-offs between batch micro-batch and pure streaming SLAs quality metrics and observability standards that make the platform trusted by every downstream the data platform in cross-team architectural forums partnering closely with ML search u0026 ranking on-device experience and platform 3. Team Leadership u0026 Management (People scope)nnPartner with recruiting to attract evaluate and hire senior and staff data engineers; raise the technical bar with every a group of data engineers directly own their performance career development and technical growth; mentor across levels on cloud-native design distributed computing and stream work against the roadmap unblock execution drive design reviews and hold a high bar for engineering craft and operational progress trade-offs and risks to senior leadership and to partner orgs; advocate for the investments the platform a healthy engineering culture: high ownership strong review practices thoughtful on-call and a deep commitment to user

MS Degree in Computer Science or related degree and 12 years of experience in Data Architecture Data Engineering or Platform Engineering with at least 5 years operating in a Principal Staff or Lead Manager experience leading and managing engineers including hiring performance management and technical mentorship of senior ICs and record of shipping petabyte-scale low-latency data platforms in production and operating them under real-world cloud expertise: expert-level proficiency with cloud object storage (e.g. AWS S3) and its architectural nuances for massive data lakes and architecting systems for entity resolution conflation or knowledge-graph construction at scale ideally involving billions of frequently updated designing pipelines that process multimodal data (structured text image) and integrate ML model inference including LLMs and embedding models: for enrichment and with LLM/model-serving infrastructure trade-offs (inference runtimes GPU-backed serving) to inform architectural decisionsnStreaming expertise: deep hands-on knowledge of Apache Kafka (or comparable brokers like Kinesis) and complex stream processing (Spark Structured Streaming Flink or similar).nData modeling: exceptional ability to design logical and physical data models for large-scale ingest retrieval and analytical consumption including dimensional modeling and lakehouse defining SLAs quality metrics and observability standards for large-scale data platforms with hands-on use of monitoring/alerting tooling (e.g. Prometheus/Grafana Datadog or OpenTelemetry-based tracing).nProgramming: command of at least one modern data-pipeline language (Scala Java or Python) and strong software engineering services integration: proven experience wiring together event notifications queuing orchestration and compute services into resilient production with vector search technologies (e.g. Pinecone Milvus) and storing/serving embeddings (e.g. pgvector Milvus FAISS) nExcellent written and verbal communication; proven ability to align engineers partner teams and senior leadership from multiple lines of business around a shared technical direction with experience bringing a consumer-oriented product from inception to production. n

Experience with embedding storage and retrieval (e.g. pgvector Milvus FAISS) and with graph databases (e.g. TigerGraph Neo4j).nExperience deploying serving and optimizing LLMs or ML models directly in the production inference runtimes/compilers (ONNX Runtime TensorRT/TensorRT-LLM) and serving frameworks (Triton vLLM TorchServe or similar). nExperience tuning batching KV-cache and GPU utilization for low-latency high-throughput real-time inference in a data pipelinenExperience with data governance tools (e.g. Apache Atlas AWS Glue Catalog DataHub).nFamiliarity with Infrastructure as Code (Terraform Pulumi) and modern CI/CD designing systems that handle petabytes of unstructured media knowledge of data privacy regulations and best practices for incorporating safety and compliance and a demonstrated instinct for building privacy-preserving systems.

Required Experience:

Manager

Were building the large-scale data foundation that powers private personalized experiences across Apple platforms. Our team designs and operates the systems that ingest unify and understand information at massive scale turning petabytes of data from many sources into a single high-quality richly st...

About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile