Data Engineer
Job Summary
This role blends DataOps / Platform Ops / Reliability Engineering for data: youll sit close to the engineering teams understand how pipelines and services behave in production and ensure we have the right tooling standards automation and runbooks to operate confidently at scale. This complements the wider engineering assurance and operational excellence expectations across the organisation.
Key Responsibilities
Platform operations & reliability
Own the operational health of the data platform: availability performance resilience and recoverability (including backup/restore and DR readiness where applicable).
Establish and maintain operational rhythms: health checks operational dashboards SLOs/SLIs alert tuning and routine maintenance.
Improve engineering hygiene across environments (dev/test/prod) ensuring consistent repeatable deployments and minimal manual steps.
Support incident management & problem elimination
Act as a primary point of contact for platform support and BAU issues triaging tickets/incidents and coordinating fixes across Data Engineering and Platform Engineering.
Lead incident response for data-platform issues (including stakeholder comms) run post-incident reviews and drive preventative actions.
Build and maintain high-quality runbooks and operational documentation.
Release & change management (data platform pipelines)
Support safe predictable releases by coordinating operational readiness: change validation rollout/rollback plans and production readiness checks.
Partner with engineers to ensure changes meet assurance standards (testing documentation observability access controls).
Observability data quality & operational governance
Alerting for critical workloads and services (including pipeline success/failure latency throughput and downstream impact).
Implement and evolve data quality checks anomaly detection and known issue playbooks to reduce repeat incidents.
Ensure platform operations align with security privacy access control and internal governance expectations.
Automation & platform evolution
Reduce operational toil through automation (scripts tooling self-service workflows CI/CD improvements).
Contribute hands-on code/config changes where needed to enable platform reliability and developer productivity (for example: deployment automation environment provisioning operational tooling).
Maintain technical awareness of new and relevant technologies and recommend improvements where they materially increase reliability security or efficiency.
Cost capacity & vendor/tooling support
Monitor platform usage and costs; identify optimisation opportunities and support capacity planning with the Data Platform Manager.
Support tooling and vendor management activities (evaluations renewals usage reporting) with pragmatic operational input.
Knowledge skills and experience required
Our core tech stack includes Python Typescript ( and React) AWS Kubernetes DataDog GraphQL PostgreSQL Mongo and Kafka.
Were looking for someone who brings:
Strong hands-on production experience operating cloud platforms and data systems (mid-level/senior individual contributor) including incident response and operational ownership.
AWS depth (e.g. building/operating data workloads and services on AWS) and comfort navigating common data-adjacent services; experience operating Glue/Lambda/Step Functions/S3/Athena-style components is a plus.
Solid engineering skills in Python (and comfort reading/modifying services or automation code) plus strong SQL.
Experience with CI/CD and infrastructure automation (e.g. Terraform and build/release tooling) with a bias for repeatability and auditability.
Practical experience with observability (monitoring/alerting/logging/tracing) ideally including DataDog and a track record of improving signal quality (not just adding alerts).
Working knowledge of containers/Kubernetes operations and the patterns required for safe production change.
Familiarity with data platform building blocks (batch streaming): orchestration ETL/ELT data modelling basics and streaming concepts (Kafka familiarity is a plus).
Strong documentation and stakeholder skills: able to explain incidents risks and trade-offs clearly to both technical and non-technical partners.
Comfort working across distributed teams proactively managing dependencies and keeping work moving without constant oversight.
Nice to have:
Experience operating data tooling such as Spark/PySpark Airflow/Prefect and/or data catalog/governance tools.
Familiarity with security controls for data platforms (secrets management IAM patterns network controls audit logging).
Working style & expectations
Youll collaborate closely with the Data Platform Manager (London-based) and the wider data engineering team (Mumbai-based) and will sometimes need working-hour overlapfor incidents releases and stakeholder communication.
You will participate in an on call/support rota (or act as escalation support) with the goal of steadily reducing incidents through automation and problem elimination.
Collinson Group is a global leader in driving loyalty and engagement for many of the worlds largest companies. Predominantly through the provision of travel related benefits within a market leading digital travel ecosystem. The group offers a unique blend of industry and sector specialists who together provide market-leading experience in delivering products and services across four core capabilities: Loyalty Lifestyle Benefits and Insurance.
The group provides unrivalled insight and expertise around affluent consumers and frequent travellers creating and delivering products and services now accessible to over 400m end consumers.
We have more than 25 years experience with 28 global locations servicing over 800 clients in 170 countries employing 1800 people.
We have been bringing innovation to the market since inception from launching the first independent global VIP lounge access Programme Priority Pass to being the first to sell direct travel insurance in the UK through Columbus Direct and creating the first loyalty agency of its kind in the travel sector with ICLP. Today we still invest heavily in innovation to ensure that we continue to deliver superior customer experiences.
Key clients include: Visa Mastercard American Express Cathay Pacific British Airways LATAM Flying Blue Accor EasyJet HSBC Chase HDFC.
Our mission is focused on doing good beyond profit which for us means we seek out opportunities for our people to share in our success and that we give back to the communities and people within which we work.
Never short of ambition the success of our business is delivered through the diverse and talented team of over 1800 colleagues globally.
Collinson is an equal opportunity employer and welcomes differences in all their forms including: colour race ethnicity gender identity sexual orientation neurodivergence family status age individuals with disabilities and people from all backgrounds cultures and experiences as we strongly believe this contributes to our on-going success.
We are focused on continually evolving our purpose driven high performing culture providing an environment where our people have the opportunity to achieve their full potential and do interesting and meaningful work. Our company values are: Act smarter Do the right thing One team and Be insight led. These help guide everything we do internally in terms of how we think act and interact right through to how we deliver value to our customers and clients.
If you need any extra support throughout the interview process then please email us at
You can look forward to a competitive salary and benefit plan including but not limited to:
Medical & Dental Care
Priority Pass Membership
Childcare allowance (for children up to 6 years old)
Birthday day off
Work From Anywhere 8 weeks per year
Required Experience:
IC