Azure Data Lead
Jersey, NJ - USA
Job Summary
Title: Azure Data Lead
Location: NJ (Day 1 Onsite)
Overview
This role is for an experienced Azure Data Lead who can design build and support scalable data engineering solutions on Microsoft Azure. The individual will work on modern data platforms involving batch and near-real-time ingestion data transformation data lake and warehouse integration and operational data workloads using Azure-native services.
The role requires strong hands-on engineering capability in PySpark Python SQL Azure Data Factory Azure Databricks Azure Data Lake Storage Azure Synapse Analytics and Azure Cosmos DB. The candidate should be able to convert business and data requirements into reliable secure performant and production-ready data pipelines.
Required skills:
- 10 years of experience in data engineering cloud data platforms ETL/ELT development or large-scale data processing
- Strong hands-on experience in designing developing testing and maintaining Azure-based data pipelines and data processing solutions
- Must have strong hands-on experience with PySpark and Python for large-scale data transformation automation data quality checks and reusable data engineering frameworks
- Azure Data Factory for data ingestion orchestration parameterized pipelines triggers and monitoring
- Azure Databricks and Apache Spark for scalable data processing using PySpark notebooks jobs workflows and optimized Spark transformations
- Azure Data Lake Storage Gen2 for lakehouse-style storage folder structures file formats access control and lifecycle management
- Azure Synapse Analytics or Azure SQL for analytical workloads SQL development data modeling performance tuning and reporting integration
- Azure Cosmos DB for NoSQL data modeling partition key design indexing strategy throughput optimization change feed processing and integration with analytics pipelines
Required technical skills:
- Strong Python programming skills including data structures functions exception handling logging reusable modules API integration and automation scripts
- Strong PySpark development experience using DataFrame APIs joins aggregations window functions UDFs partitioning caching broadcast joins and performance optimization
- Good SQL skills for querying transformation data validation stored procedures performance tuning and troubleshooting data issues
- Experience with Git Azure DevOps CI/CD practices unit testing deployment pipelines monitoring and production support for data engineering workloads
Responsibilities:
Data Engineering Design & Development
Design develop and maintain scalable Azure data engineering solutions across:
- Batch incremental and near-real-time data ingestion from databases APIs files applications and streaming sources
- Azure Data Factory Azure Databricks ADLS Gen2 Azure Synapse Analytics Azure SQL and Azure Cosmos DB
Build and optimize data pipelines for:
- Data extraction cleansing transformation enrichment validation and loading into curated data layers
- Reusable PySpark frameworks parameterized notebooks modular Python components and metadata-driven processing patterns
- Data quality controls exception handling audit logging reconciliation restartability and operational monitoring
Develop Cosmos DB-based data solutions by:
- Designing containers partition keys indexing policies consistency levels TTL and throughput configuration based on access patterns
- Implementing ingestion and integration patterns between Cosmos DB Azure Data Factory Databricks ADLS and analytical stores
- Using Cosmos DB change feed bulk operations query tuning partition-aware design and cost optimization practices
Pipeline Delivery Optimization & Support
Own hands-on delivery across:
- PySpark-based ETL/ELT jobs for large-scale structured semi-structured and unstructured data processing
- Python-based automation data validation utilities reusable transformation logic and integration scripts
- Azure Data Factory pipelines Databricks jobs Synapse SQL workloads Cosmos DB integrations and downstream analytics data products
Drive engineering discipline through:
- Code reviews unit testing version control CI/CD deployment automation and environment configuration management
- Pipeline monitoring failure handling performance tuning cost optimization and production incident resolution
Preferred Qualifications:
- Microsoft Azure Data Engineer certification or equivalent hands-on Azure project experience
- Experience with Delta Lake lakehouse patterns medallion architecture and data warehouse modeling
- Exposure to event-driven or streaming patterns using Event Hubs Kafka Stream Analytics or Databricks Structured Streaming
- Understanding of data security RBAC managed identities private endpoints encryption and compliance-driven data handling
Key Attributes:
- Strong analytical and problem-solving skills with the ability to troubleshoot complex data and pipeline issues
- Ability to work with business analysts architects QA teams and client stakeholders to clarify requirements and deliver reliable data solutions
- Good communication skills with the ability to explain technical designs pipeline behavior and production issues clearly
- Ownership mindset with focus on quality maintainability performance security and operational stability
Must Have skills:
- Cosmos DB
- Azure Data Lake Storage
- Azure Synapse Analytics
- Azure Data Factory
- Postgre SQL
Regards
Manoj Goud
Derex Technologies INC
Contact : Ext 206
Additional Information :
All your information will be kept confidential according to EEO guidelines.
Remote Work :
No
Employment Type :
Full-time
About Company
Derex Technologies Inc specializes in providing IT consulting, staffing solutions and software services. Globally headquartered in Harrison New Jersey since 1996 Derex delivers the highest quality technology professionals and an array of customized IT talent solutions designed to impr ... View more