Data Engineer ADAS Fleet Analytics
San Jose, CA - USA
Job Summary
- Design and maintain scalable telemetry pipelines using Event Hub Medallion architecture Delta Lake MERGE schema evolution and blue/green data deployments.
- Own campaign and fleet management including enrollment rosters ingestion-completeness monitoring VIN reconciliation and automated coverage alerts.
- Develop fleet KPIs for engineering including safety events and takeover analysis; operations including data freshness and pipeline health; and management including trends and cohort comparisons.
- Build ML-ready data infrastructure including video and signal pipelines for SSR recordings embedding stores feature-engineering layers and model-output integration into Gold.
- Operate the Azure data platform including ADLS Gen2 Synapse Spark Event Hub and Container Apps and support Terraform-managed infrastructure and CI/CD.
- Write automated tests maintain documentation alongside code and participate in code reviews and incident response.
- Bachelors or Masters degree in Computer Science Data Engineering or a related field. Equivalent experience may be considered.
- 2-5 years of experience in data engineering or big-data analytics with production pipeline ownership.
- Strong Python PySpark DataFrame API and performance-tuning skills and SQL skills including window functions and aggregations.
- Experience with a columnar or transactional data-lake technology such as Delta Lake Iceberg Hudi or managed Parquet.
- Experience with automated testing for data pipelines and Git-based development workflows.
- Ability to diagnose production failures including schema conflicts data skew and memory issues and communicate trade-offs to technical and non-technical stakeholders.
Strongly Preferred
Medallion or other layered data-processing patterns.
Cloud data-platform experience in Azure AWS or GCP.
Parquet schema design and schema evolution.
Production data-quality practices including deduplication idempotency and data contracts.
Time-series IoT or vehicle-telemetry data.
Nice to Have
DuckDB or PyArrow; FastAPI or analytical API delivery; blue/green data deployments.
ML data infrastructure including feature stores embedding pipelines vector databases or video pipelines.
Fleet or campaign management for vehicle or IoT programs.
Automotive data privacy including CCPA; ADAS or connected-vehicle domain knowledge.
Azure Data Engineer or Databricks certification.
English required; German is an advantage.