Enter a job title or keyword

Data Engineer (DatabricksAWS)


Job Location:

Indianapolis, IN - USA

Monthly Salary: Not provided by the employer
Experience Required: 5years
Posted: 7 August 2026 (30+ days ago)
Application Deadline: 4 November 2026
Vacancies: 1 Vacancy

Job Summary

AWS Data Engineer - Databricks
Hybrid Indianapolis IN

About the Role

We are seeking a Data Engineer with 35 years of experience working specifically within the pharma industry to join a pharma-focused data team. This is a senior-flavored engineering role that combines hands-on pipeline and platform work with significant business-facing responsibility including translating business needs into technical specs presenting to executive-level stakeholders and helping stand up new data domains from the ground up. You will design build and govern the data infrastructure that powers analytics and reporting across the business while also acting as a trusted technical partner to non-technical stakeholders.

Key Responsibilities
  • Design build and maintain scalable ETL/ELT pipelines (batch and streaming) using Databricks AWS and related orchestration tools
  • Write and optimize advanced SQL and build data transformations in Python or Scala
  • Integrate external data sources via APIs and manage pipeline orchestration (Airflow Databricks Workflows AWS Glue)
  • Apply data quality governance cataloging and lineage practices aligned with regulated-industry standards
  • Work within GxP-regulated data environments and apply awareness of data privacy/compliance considerations (e.g. 21 CFR Part 11 GDPR where applicable)
  • Partner with business stakeholders across the pharma value chain (R&D Manufacturing & Quality Commercial Drug Development) to gather and translate requirements into technical specifications
  • Present technical work and data strategy to executive-level audiences
  • Prioritize high-impact data initiatives and proactively identify and avoid duplicated data efforts
  • Support change management and adoption of new data solutions across business teams
  • Help stand up new data domains from scratch (green-field build) not just maintain existing ones


Requirements
Required Qualifications
Data Engineering & Pipelines
  • ETL/ELT development (batch and streaming)
  • Advanced SQL (joins window functions query optimization)
  • Python or Scala for data transformation
  • Data pipeline orchestration (Airflow Databricks Workflows AWS Glue)
  • API integration for external data source ingestion
Platforms & Tools
  • Databricks (Delta Lake Unity Catalog Genie)
  • Cloud platforms AWS (S3 Glue Athena) and/or Azure/GCP equivalents
  • Data warehousing concepts (dimensional modeling star schema)
  • BI/visualization tools (Tableau Power BI or similar) to understand downstream consumption
Data Quality & Governance
  • Data profiling and cleansing techniques
  • Metadata management and data cataloging
  • Master data management (MDM) principles
  • Data lineage tracking
  • Data governance frameworks (especially regulated-industry standards)
Pharma / Life Sciences Domain Knowledge
  • Familiarity with GxP-regulated data environments
  • Understanding of the pharma value chain (R&D Manufacturing & Quality Commercial Drug Development)
  • Awareness of data privacy/compliance considerations (21 CFR Part 11 GDPR where applicable)
  • Knowledge of common pharma data domains (clinical manufacturing quality commercial)
Stakeholder Management
  • Requirements gathering and translation (business need technical spec)
  • Cross-functional communication (Business IT)
  • Executive-level presentation skills (given EC visibility)
  • Change management / adoption support
Analytical & Strategic Thinking
  • Prioritization frameworks (identifying high-impact vs. low-value data asks)
  • Cost-avoidance mindset (spotting duplication before it happens)
  • Ability to work with ambiguity and evolving priorities
Project & Program Skills
  • Agile/Scrum familiarity
  • Documentation discipline (data dictionaries source-to-target mappings)
  • Vendor/partner coordination (if external data sources are involved)
Nice-to-Have Differentiators
  • Prior consulting or client-facing delivery experience
  • Experience standing up new data domains from scratch (green-field vs. maintenance)
  • Familiarity with AI/GenAI-enabled analytics tools


Required Skills:

Required Qualifications 35 years of data engineering experience specifically within the pharma industry Heavy hands-on experience with Databricks Heavy hands-on experience with AWS data services (e.g. S3 Glue Redshift Lambda Kinesis) Strong experience with Big Data technologies and Apache Spark (PySpark preferred) Demonstrated experience building and maintaining ETL/ELT pipelines end-to-end Experience with orchestration tools (e.g. Apache Airflow) and CI/CD practices Experience with monitoring/observability for data pipelines Demonstrated ability to understand pharma business needs and speak to pharma business groups Strong SQL and Python skills Bachelors degree in computer science Data Engineering or a related field or equivalent practical experience Preferred Qualifications Databricks Data Engineer certification Experience with Delta Lake Snowflake or similar modern data platforms Experience with Infrastructure-as-Code and containerization (Docker Kubernetes) Prior experience supporting pharma commercial clinical or R&D data functions