Data Engineer (DatabricksAWS)
Indianapolis, IN - USA
Job Summary
- Design build and maintain scalable ETL/ELT pipelines (batch and streaming) using Databricks AWS and related orchestration tools
- Write and optimize advanced SQL and build data transformations in Python or Scala
- Integrate external data sources via APIs and manage pipeline orchestration (Airflow Databricks Workflows AWS Glue)
- Apply data quality governance cataloging and lineage practices aligned with regulated-industry standards
- Work within GxP-regulated data environments and apply awareness of data privacy/compliance considerations (e.g. 21 CFR Part 11 GDPR where applicable)
- Partner with business stakeholders across the pharma value chain (R&D Manufacturing & Quality Commercial Drug Development) to gather and translate requirements into technical specifications
- Present technical work and data strategy to executive-level audiences
- Prioritize high-impact data initiatives and proactively identify and avoid duplicated data efforts
- Support change management and adoption of new data solutions across business teams
- Help stand up new data domains from scratch (green-field build) not just maintain existing ones
- ETL/ELT development (batch and streaming)
- Advanced SQL (joins window functions query optimization)
- Python or Scala for data transformation
- Data pipeline orchestration (Airflow Databricks Workflows AWS Glue)
- API integration for external data source ingestion
- Databricks (Delta Lake Unity Catalog Genie)
- Cloud platforms AWS (S3 Glue Athena) and/or Azure/GCP equivalents
- Data warehousing concepts (dimensional modeling star schema)
- BI/visualization tools (Tableau Power BI or similar) to understand downstream consumption
- Data profiling and cleansing techniques
- Metadata management and data cataloging
- Master data management (MDM) principles
- Data lineage tracking
- Data governance frameworks (especially regulated-industry standards)
- Familiarity with GxP-regulated data environments
- Understanding of the pharma value chain (R&D Manufacturing & Quality Commercial Drug Development)
- Awareness of data privacy/compliance considerations (21 CFR Part 11 GDPR where applicable)
- Knowledge of common pharma data domains (clinical manufacturing quality commercial)
- Requirements gathering and translation (business need technical spec)
- Cross-functional communication (Business IT)
- Executive-level presentation skills (given EC visibility)
- Change management / adoption support
- Prioritization frameworks (identifying high-impact vs. low-value data asks)
- Cost-avoidance mindset (spotting duplication before it happens)
- Ability to work with ambiguity and evolving priorities
- Agile/Scrum familiarity
- Documentation discipline (data dictionaries source-to-target mappings)
- Vendor/partner coordination (if external data sources are involved)
- Prior consulting or client-facing delivery experience
- Experience standing up new data domains from scratch (green-field vs. maintenance)
- Familiarity with AI/GenAI-enabled analytics tools
Required Skills:
Required Qualifications 35 years of data engineering experience specifically within the pharma industry Heavy hands-on experience with Databricks Heavy hands-on experience with AWS data services (e.g. S3 Glue Redshift Lambda Kinesis) Strong experience with Big Data technologies and Apache Spark (PySpark preferred) Demonstrated experience building and maintaining ETL/ELT pipelines end-to-end Experience with orchestration tools (e.g. Apache Airflow) and CI/CD practices Experience with monitoring/observability for data pipelines Demonstrated ability to understand pharma business needs and speak to pharma business groups Strong SQL and Python skills Bachelors degree in computer science Data Engineering or a related field or equivalent practical experience Preferred Qualifications Databricks Data Engineer certification Experience with Delta Lake Snowflake or similar modern data platforms Experience with Infrastructure-as-Code and containerization (Docker Kubernetes) Prior experience supporting pharma commercial clinical or R&D data functions