Foundation Engineering SRE Platforms Site Reliability Engineer – Associate London
Job Summary
Role Overview
Goldman Sachs has embarked on one of its most ambitious engineering programs: theConsolidated Trade Ledger (CTL) a ground-up reimagining of the front-to-back architecture that underpins every trade the firm executes. CTL is a flagship initiative jointly sponsored by Global Markets and Engineering leadership and it sits at the very heart of the firms core technology strategy. This new cloud-native platform will deliver the capacity extensibility scalability and innovation capabilities to power the next two decades of growth for Global Markets while simultaneously driving significant operational efficiencies.
We are seeking a Site Reliability Engineer to help build run and continuously improve this business-critical service. The role combines software engineering systems engineering and production expertise to improve the reliability scalability observability incident response and operational efficiency of the CTL platform.
Required Skills / Experience
- Strong programming ability in one or more modern languages Java or Go strongly preferred with experience building maintainable automation beyond simple scripts.
- Good understanding of networking messaging distributed systems data structures algorithms and software design fundamentals.
- Hands-on experience with observability tooling including metrics logging tracing and dashboarding platforms such as Prometheus Grafana ELK or OpenTelemetry.
- Proven ability to investigate production issues identify root causes and deliver durable engineering fixes that improve system behaviour and reduce repeat incidents.
- Excellent written and verbal communication skills with the ability to translate complex technical issues into clear updates for technical business and senior stakeholders.
Preferred Qualifications / Skills / Experience
- Bachelors degree in Computer Science Engineering or a related technical field or equivalent practical experience.
- Experience with public cloud platforms GCP preferred cloud-native architecture Kubernetes microservices or service mesh technologies.
- Familiarity with relational databases and/or data-intensive platforms
- Experience supporting mission-critical production systems preferably in a financial services environment.
- Experience implementing progressive delivery approaches such as canary releases blue/green deployments feature flags or chaos/game days.
- Experience developing AI-assisted operations capabilities such as alert enrichment anomaly detection triage support or runbook automation.
Key SRE Competencies
- Reliability engineering: SLOs SLIs error budgets incident reduction and service health.
- Production excellence: monitoring alerting observability runbooks handoffs and operational readiness.
- Engineering mindset: automation coding debugging testing scalable design and systems thinking.
- Incident leadership: calm execution escalation discipline clear communication and blameless learning.
Required Experience:
IC
About Company
The Goldman Sachs Group, Inc. is a leading global investment banking, securities, and asset and wealth management firm that provides a wide range of financial services.