Applied AI & Data Engineer Business & Education
Cupertino, CA - USA
Job Summary
Apples Business and Education organization builds the infrastructure platforms and services behind Apples offerings for enterprise and education customers: device management identity and subscription services and the classroom apps built on top of them. Our team owns the data engineering behind all of it from pipelines and lakehouse architecture through analytics reporting and we are building a new generation of AI-native capabilities on top: agents intelligent workflows and self-serve analytics that change how our Data Engineering Analytics and Data Science teams work. The ambition is to define what an AI-first data organization looks like at Apple are looking for a well-rounded builder. You spent the earlier part of your career deep in software data or ML engineering and the last few years applying that foundation to ship Applied AI products end-to-end. You think architecturally you know where LLMs and agents earn their keep and where deterministic code is the better answer and you would rather measure a systems quality than argue about it.
Design and build the data and AI platform for Business and Education on Databricks AWS and modern cloud-native patterns along with the data models and data-quality practices that make it AI and agentic systems end-to-end: retrieval planning evaluation guardrails responsible-AI review deployment and the on-call rotation that keeps them the data foundation for GenAI agentic AI and advanced analytics: RAG pipelines vector search knowledge graphs and multi-agent orchestration so the organization can ship natural-language data interfaces AI agents tool-calling workflows and data-driven web those systems efficient enough to scale through model selection and serving decisions latency and throughput work and token economics at Apple with product business analytics and AI stakeholders to turn ambiguous requirements into secure scalable production-ready hands-on technical leadership through design reviews implementation guidance and production-readiness checks and own projects across their full lifecycle from discovery and planning through engineers prioritize and resource across concurrent initiatives and help the team adopt AI-native practices as they emerge through workshops technical playbooks and design state-of-the-art data and AI techniques including agentic patterns evaluation methods AI-native developer tools and modern data architectures and turn them into capabilities that make our Data Engineering Analytics and Data Science teams measurably faster: AI-accelerated pipeline development intelligent alerting and natural-language access to data.
8 years across data engineering analytics engineering software engineering or ML engineering with the last 3 years building and shipping Applied AI and agentic LLM systems in production. You are still a builder: you want to spend real time writing code prototyping and shipping alongside your team not only reviewing what others architect build and operate production AI products composed of LLMs foundation models agents and deterministic components for both human and machine consumers. You have clear judgment on where to infer and where to compute how to decompose tasks across specialized models how to orchestrate multi-step reasoning and tool use and how the system degrades when a model -on fluency with modern LLM and agent frameworks (LangChain LlamaIndex Semantic Kernel Google ADK or equivalent) vector search (pgvector FAISS Pinecone or equivalent) RAG pipelines multi-agent coordination tool invocation and stateful reasoning. You have moved past vanilla RAG: you know where retrieval breaks and when to reach for planning reranking structured reasoning fine-tuning or plain deterministic code discipline for AI systems: evaluation harnesses guardrails and telemetry that change decisions (offline evals golden sets LLM-as-judge behavioral regression drift monitoring) and optimization for cost latency throughput and inference quality (model selection serving decisions token-spend control caching batching streaming distillation quantization speculative decoding).nA foundation in machine learning and deep learning. You understand how transformers and LLMs are trained fine-tuned and evaluated you reason about embeddings loss functions and statistical rigor and you can tell whether a production failure is prompt retrieval model or design and build scalable data platforms on modern cloud-native patterns (Databricks AWS or equivalent) and you are as comfortable in the warehouse and the SQL engine (Trino Presto Spark) as in the model-serving layer. Proficiency in at least one high-level language (Python Scala Java or Go) strong SQL and the discipline to write code that is readable observable in production and testable at the delivering ETL/ELT streaming and CDC (change data capture) pipelines with technologies such as Spark Kafka and Delta Lake for both batch and real-time data along with the workflow orchestration data quality checks observability and alerting that catch breakage before it reaches downstream analytics or AI of data modeling patterns and the judgment to pick the right one for a given use case trading off analytical query performance governance and track record of hands-on technical leadership: architecture and design reviews implementation guidance production-readiness review and 3 years mentoring engineers and prioritizing across concurrent initiatives. You communicate clearly enough across cross-functional teams to influence strategy and you raise the AI fluency of partner organizations through workshops playbooks and design product mindset paired with a research sensibility. You read papers separate signal from hype work loosely defined problems with meticulous attention to detail and drive them to completion without sacrificing trust in the or MS in Computer Science Information Systems Artificial Intelligence Machine Learning Engineering Mathematics Statistics or a related field or equivalent practical experience building data and AI systems in production.
Model and prompt customization at scale: fine-tuning foundation models training reward models building custom retrieval reranking or embedding models for domain-specific tasks and prompt engineering optimized for performance reliability and with MLOps and LLMOps: model lifecycle management deployment pipelines observability and prompt and evaluation building natural-language interfaces over data text-to-SQL semantic search or analytics copilots for internal or customer-facing using AI-native code editors and agent-assisted development environments to improve developer productivity and establishing guardrails for their responsible use across security IP protection compliance and code with Google Cloud or Azure stream-processing systems (Apache Flink Spark Streaming Kafka Streams) and NoSQL or analytics datastores (Cassandra MongoDB Druid Apache Pinot) for real-time data and real-time AI building AI machine learning and experimentation systems in regulated or privacy-sensitive to open source research talks or technical writing that have shaped how others build AI experience leading or managing engineers or serving as technical lead across multiple concurrent data and AI projects.
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more