Data Engineer
Job Summary
CodiLime is a software and network engineering industry expert and the first-choice service partner for top global networking hardware providers software providers and telecoms. We create proofs-of-concept help our clients build new products nurture existing ones and provide services in production environments. Our clients include both tech startups and big players in various industries and geographic locations (US Japan Israel Europe).
While no longer a startup - we have 250 people on board and have been operating since 2011 weve kept our people-oriented culture. Our values are simple:
Act to deliver.
Disrupt to grow.
Team up to win.
You will join the team behind a large-scale centralized data platform built for a global consulting organization. The platform is the shared source of company data behind several of the firms internal products and is used daily by consultants for company research including support for Mergers & Acquisitions (M&A) engagements.
This is fundamentally a data engineering role combined with software engineering: youll be designing coding testing and operating production Python systems - pipelines libraries and services - that move transform and serve data at scale. This is not a role focused on configuring tools or writing one-off queries. Youll be building reliable maintainable software that powers our data platform. The goal is a unified enterprise-grade dataset of 300M company records integrated from 10 external and internal sources.
The platform delivers firm-level and site-level data - firmographics technographics and hierarchical relationships (parent company subsidiary site) - alongside key business metrics such as revenue CAGR EBITDA headcount M&A activity competitors industry classification and web traffic. Data needs to stay accurate well-structured and fast to query as both the dataset and the number of consumers keep growing.
Technology stack:
Languages: Python SQL
Data platform: Snowflake dbt
Workflow orchestration: Apache Airflow (complex DAGs) running on Kubernetes
Data processing: Apache Spark on Azure Databricks
Data tooling and DBs: pandas Polars PyArrow DuckDB PySpark PostgreSQL Redis
Cloud: Azure (AKS Blob Storage ACR Databricks OpenSearch Azure AI Search)
API & services: FastAPI (REST async) API Gateway
Testing & code quality: pytest mypy/pyright ruff/black sqlfluff SonarQube
Schema validation: Pydantic
Dependency & environment management: uv Poetry
CI/CD & infrastructure: GitHub Actions Docker Kubernetes
AI-Assisted Development: Cursor Claude Code ChatGPT Enterprise
Future direction: agentic AI systems LangChain Azure OpenAI integration
What else you should know:
Team: Data Architecture Lead Data Engineers DataOps Engineers Backend Engineer Product Owner collaboration with Frontend Engineers and Data Science and AI Engineers
Distributed team across Europe and India
Agile collaborative environment; given the platforms organization-wide impact were looking for a mature proactive results-driven approach
Code quality is enforced through testing typing and tooling - not just code review
We work on multiple interesting projects at a time so it may happen that well invite you to an interview for another project if we see that your competencies and profile are well suited for it.
This is a results-driven contributor role split roughly between hands-on engineering delivery and continuous improvement of the platform.
As a part of the project team you will be responsible for:
Engineering & delivery (70%)
Design build and maintain batch and streaming data pipelines in Python including their orchestration scheduling and monitoring (Airflow)
Write reusable well-typed Python libraries and internal packages used by other engineers and analysts
Build Python services and APIs (FastAPI) that expose data to downstream applications and integrate with third-party and internal APIs
Develop data transformation logic in SQL and dbt on Snowflake and build data models that support fast reliable querying
Write unit integration and data-contract tests and keep pipelines covered by automated CI
Profile and optimize Python code and data processing jobs for runtime memory and cost
Deploy and operate code in the cloud using containers infrastructure as code and CI/CD
Enforce access controls secrets handling and sensitive-data protections throughout the data lifecycle
Using AI coding assistants effectively while validating all generated output before it reaches production and helping build automated quality gates
Continuous improvement (30%)
Find and fix efficiency reliability cost and correctness issues in existing pipelines refactoring toward simpler designs
Replace one-off scripts and notebooks with tested packaged scheduled code
Implement data quality checks validation and monitoring that catch issues before consumers do
Create matching logic to deduplicate and connect entities across multiple data sources
Document data processes and system architecture and maintain project documentation
As a Data Engineer you must meet the following criteria:
Strong Python experience: data structures typing error handling generators/iterators context managers and the standard library
Software engineering fundamentals: modular design dependency management packaging and API design
Testing discipline: pytest fixtures mocking and writing code thats testable by construction
Hands-on experience building and operating ETL/ELT pipelines in production not just scripts or notebooks
Strong experience with Snowflake and dbt
Experience with Apache Airflow or similar code-based orchestration tools
Solid working knowledge of SQL and data modeling sufficient to design robust database schemas and query them effectively
Experience with Docker Kubernetes and CI/CD practices
Debugging and profiling skills; able to reason about performance concurrency and memory in Python
Experience with at least one public cloud (AWS or Azure)
Experience with version control systems (Git)
Able to explain technical trade-offs to both technical and business audiences
Experience using AI coding assistants such as Claude Code Cursor or similar on a daily basis
Strong communication skills and good knowledge of English (minimum C1 level)
Beyond the criteria above we would appreciate the following nice-to-haves:
Experience with Apache Spark ideally on Databricks
Experience with Pydantic or similar schema-validation libraries
Python web/API frameworks (FastAPI Flask)
Async Python multiprocessing or other concurrency patterns
Experience with Azure AI Search or AWS OpenSearch
A second language: Go Rust Scala or TypeScript
Familiarity with LLMs Azure OpenAI or agentic AI systems
Flexible working hours and approach to work: fully remotely in the office or hybrid
Professional growth supported by internal training sessions and a training budget
Solid onboarding with a hands-on approach to give you an easy start
A great atmosphere among professionals who are passionate about their work
The ability to change the project you work on