Data Governance Engineer
Austin, TX - USA
Job Summary
At Apple Ads we are building the next generation of privacy-focused advertising the Data Governance team we work at the cutting edge of data engineering machine learning and privacy at Apples scale. We are constantly developing data and privacy management products to provide amazing user experiences and to drive value for developers and team is seeking a Data Governance Engineer to support our data governance objectives. You will play a crucial role in delivering on Apples privacy commitments to our customers. You will partner across our engineering product privacy and reliability organizations to deliver on data access use protection minimization retention and other data governance execution areas.
Design autonomous governance agents for petabyte-scale data pipelines (e.g. Kafka Spark Hadoop) ensuring compliance and quality across all processing stagesnBuild MCP integrations enabling automated policy enforcement PII detection and data quality validation across diverse storage and compute environmentsnPartner with engineers and PMs to develop intelligent monitoring that detects anomalies generates compliance reports and auto-remediates issuesnCreate reusable agent templates and workflows that adapt to changing regulations while maintaining performance and reducing manual compliancenContribute to distributed agent ecosystem emphasizing reliability through self-healing mechanisms automated failovers and dynamic resource allocation for scalabilitynUse AI coding agents to accelerate implementation test coverage and debugging validating every result for correctness privacy and cost
B.S. in Computer Science or related field with 5 years of software development experience including exposure to ML/AI applicationsnProduction experience in Python Java Rust or similar languages with familiarity in data processing and API developmentnUnderstanding of distributed systems and cloud platforms including CAP theorem tradeoffs and basic ML model deploymentnExperience with containerization (Docker) orchestration (Kubernetes) and infrastructure as code (e.g. Terraform CloudFormation)nProficiency in CI/CD pipelines and DevOps practices using Git GitHub Actions/Jenkins/GitLab CI with experience in automated testing and deployment workflowsnFamiliarity with observability and monitoring tools (e.g. Prometheus Grafana DataDog) and logging frameworks for production systemsnBasic knowledge of machine learning concepts and MLOps including data pipelines model versioning and experiment tracking tools
Expertise in open source data analytics and governance platforms: architecture deployment and performance tuning of Datahub Apache Spark Flink Hive Hadoop/HDFS and Iceberg Rest CatalognExperience building multi-agent AI systems: proficiency with LangChain LangGraph or AutoGen frameworks; strong prompt engineering and LLM integration skills; ability to design event-driven architectures for autonomous workflowsnSkills in integration and communication layers: implement MCP servers and APIs using Python REST/GraphQL and message queuing (e.g. Kafka RabbitMQ); experience with modern data platforms including Snowflake Databricks and vector databasesnMLOps and observability capabilities: deploy containerized AI systems with comprehensive monitoring; track experiments using MLflow or Weights u0026 Biases; implement distributed tracing for agent workflows and model performance
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more