Software Engineer
Chicago, IL - USA
Job Summary
MalaceHR is seeking an experienced Data Engineer to join a fast-paced technology and enterprise data environment. This position will be responsible for designing developing and optimizing scalable data solutions using modern cloud and big data technologies.
The ideal candidate is a self-starter with strong hands-on experience in Azure Databricks Python PySpark Spark SQL Azure Functions Delta Lake Azure DevOps CI/CD and agent-driven workflows leveraging MCP frameworks. This individual will work closely with cross-functional technical and business teams to develop data pipelines modernize data platforms and support enterprise-level data initiatives.
- Design architect and develop scalable data solutions using cloud-based big data technologies.
- Ingest process transform and analyze large and diverse datasets to support business and technical requirements.
- Design and develop data management and persistence solutions using relational and non-relational databases.
- Build scalable data lake solutions capable of storing structured semi-structured and unstructured data from internal and external sources.
- Develop proofs of concept (POCs) to validate new technologies architectures and solution proposals.
- Provide technical guidance supporting migrations to modern cloud-based data platforms.
- Develop maintain and optimize ETL and ELT workflows using Azure Databricks Python PySpark and Spark SQL.
- Extract manipulate and transform data from databases data lakes APIs files streaming platforms and other data sources.
- Develop efficient data-processing pipelines with appropriate validation error handling monitoring and performance tuning.
- Build systems that ingest cleanse normalize and structure large and diverse datasets.
- Develop event-driven and streaming data pipelines supporting real-time and near-real-time processing.
- Contribute to and follow CI/CD processes and Data Engineering development best practices.
- Perform unit testing system integration testing regression testing and support user acceptance testing.
- Translate business requirements into scalable technical solutions that can be designed and engineered.
- Partner with business stakeholders to develop technical documentation and communication materials supporting proper data usage and interpretation.
- Implement data security best practices including encryption access controls data governance and compliance requirements.
- Maintain data privacy confidentiality quality and integrity throughout the data engineering lifecycle.
- Perform data analysis and troubleshooting to identify and resolve data-related issues.
- Build agent-driven workflows used to load process and transform data.
- Develop data solutions leveraging MCP frameworks and agent orchestration.
- Collaborate with software engineers architects product teams analysts and other cross-functional stakeholders.
- Minimum of 5 years of professional Data Engineering or Data Development experience.
- Strong hands-on experience with:
- Python
- PySpark
- Spark SQL
- ETL/ELT development
- SQL Server
- Data pipeline development
- Bachelors degree in Computer Science Information Science Mathematics Statistics Engineering Data Science or another related quantitative discipline preferred.
- Experience working with Microsoft Azure cloud technologies.
- Strong experience with Azure Databricks and Azure Storage/Data Lake technologies.
- Strong understanding of data engineering architecture data modeling data integration and ETL concepts.
- Excellent analytical technical troubleshooting and organizational skills.
- Strong written and verbal communication skills including technical documentation.
- Hands-on experience working with structured semi-structured and unstructured data.
- Experience developing solutions within data lake environments.
- Experience building event-driven data pipelines utilizing queues streaming platforms or similar technologies.
- Strong understanding of real-time and near-real-time data processing.
- Advanced hands-on experience with PySpark Databricks and Spark SQL.
- Experience working with file and serialization formats including:
- JSON
- Parquet
- CSV
- Other structured and semi-structured data formats
- Knowledge of NoSQL database technologies such as:
- Cosmos DB
- MongoDB
- HBase
- Similar distributed database platforms
- Experience with cloud environments such as Microsoft Azure or AWS with Azure preferred.
- Experience with technologies including:
- Azure SQL Server
- Azure Data Lake Storage
- Azure Event Hubs
- Azure Functions
- Azure Search
- Cosmos DB
- MongoDB
- Spark Streaming
- Delta Lake
- Azure DevOps
- CI/CD pipelines
- Experience developing or working with AI agent orchestration and MCP-based workflows is highly desirable.
- Ability to troubleshoot complex data-processing and pipeline performance issues.
- Ability to manage multiple projects and priorities simultaneously in a fast-paced environment.
- Reliable self-motivated organized and capable of working independently.
- Strong team player with the ability to collaborate effectively with cross-functional technical and business teams.