Mids-Level Data Engineer
Cape Town - South Africa
Job Summary
Core experience
- 35 years professional experience in Data Engineering data-focused backend engineering or data architecture.
- Proven commercial experience actually building maintaining and troubleshooting production ETL/ELT pipelines rather than purely academic/project exposure.
- Experience working with data warehouses data architecture data modelling and production data environments.
- Bachelors degree in Computer Science Software Engineering Data Engineering Information Systems or related field or equivalent practical experience.
Python & SQL essential
- Strong/advanced Python for data engineering ideally including Pandas and PySpark; FastAPI exposure is also specified at mid-level.
- Strong to expert SQL including complex joins window functions debugging query optimisation and query-plan/performance optimisation.
- Must be capable of troubleshooting and optimising both SQL queries and Python data jobs.
- This is a significant upgrade from the original requirement where only good SQL and basic programming/scripting were required.
ETL/ELT & pipelines essential
- Design build test maintain and optimise scalable ETL/ELT pipelines.
- Batch processing experience with real-time/streaming exposure highly valuable.
- Data ingestion from multiple sources including:
- REST APIs
- PostgreSQL/MySQL or other relational databases
- Third-party/SaaS platforms
- Structured semi-structured and unstructured data
- Ideally Kafka/RabbitMQ or other message queues.
- Production pipeline monitoring error handling debugging incident/root-cause resolution and preventative improvements.
Data warehousing & modelling essential
- Hands-on experience with modern data warehouse/lakehouse platforms such as Snowflake BigQuery Amazon Redshift or Databricks.
- Strong understanding of data models schemas and dimensional modelling.
- Practical exposure to Kimball/star schema; Data Vault is advantageous.
- Understanding of storage/query optimisation including indexing partitioning and compression.
Cloud essential for the upgraded role
- Solid working knowledge of at least one major cloud platform:
AWS Azure or GCP. - Hands-on exposure to the platforms data services rather than simply having a cloud certification.
- The junior specification only required familiarity with cloud platforms such as Fabric AWS or BigQuery; the upgraded role requires solid working knowledge.
Orchestration & modern data stack essential
- Commercial experience with Apache Airflow dbt Prefect or a comparable orchestration/workflow framework.
- Experience transforming data using dbt Spark or cloud-native warehouse tools.
- Understanding of modern scalable data architecture.
Git & CI/CD
- Strong Git experience.
- GitHub/GitLab workflows.
- Basic CI/CD pipeline automation for deploying data code.
- Should be accustomed to collaborative development code reviews and version-controlled production code.
Data quality governance & security
- Automated data validation/testing preferably dbt tests Great Expectations or equivalent.
- Data quality monitoring and alerting.
- Data lineage metadata and data dictionaries.
- Understanding of access controls masking and encryption.
- Awareness of data governance/compliance requirements.
AI/ML data engineering
- Must understand how Data Engineering supports AI/ML workloads.
- Preparing clean reproducible datasets for model training evaluation and inference.
- Data preprocessing and validation for ML pipelines.
- Exposure to feature stores MLOps and generative-AI integrations is highly advantageous.
- The upgraded requirements build directly on the newer Junior specification which already introduced structured/unstructured data preparation vector databases and RAG pipelines.
- Bonus technologies include MLflow Feast Vertex AI and SageMaker.
Highly advantageous / differentiators
- Kafka AWS Kinesis or Apache Flink
- Docker
- Kubernetes
- Terraform or CloudFormation
- NoSQL MongoDB DynamoDB or Redis
- Spark/PySpark
- Real-time streaming pipelines
- MLOps
- Cost optimisation within cloud data environments.
35 years relevant commercial experience strong Python advanced SQL hands-on ETL/ELT pipeline development data warehousing / modelling AWS / Azure / GCP Airflow / dbt / Prefect or equivalent Git genuine production troubleshooting / optimisation experience production Data Engineering experience.
additional advantage Snowflake / Databricks / BigQuery / Redshift PySpark/ Spark CI/CD automated data testing AI/ML data pipelines and ideally Kafka/streaming.