Sovereign Engineering Platform SRE (mfd)
Job Summary
Key responsibilities
- Build and operate Kubernetes environments that host AI engineering tools internal model gateways retrieval components workflow services CI/CD runners and documentation services.
- Implement GitOps and Infrastructure as Code patterns for reproducible provisioning configuration policy enforcement platform upgrades and disaster recovery readiness.
- Manage private registries package mirrors secrets identity integration network segmentation storage classes backup routines and controlled connectivity models.
- Provide observability for engineering workloads including metrics logs traces GPU and CPU utilization service health cost signals and operational runbooks.
- Work with software security and architecture teams to ensure the platform supports AI-assisted SDLC workflows without creating uncontrolled data exposure or audit gaps.
Examples of market tools models and platform components expected
- Platform tooling such as Kubernetes Helm Terraform Ansible ArgoCD Crossplane GitLab runners Jenkins agents private registries and internal package mirrors.
- AI platform components such as vLLM Ollama OpenAI-compatible gateways Qdrant or similar vector stores Open WebUI Continue-compatible endpoints and workflow services.
- Observability and operations stacks such as Prometheus Grafana Loki OpenTelemetry ELK/OpenSearch Alertmanager SRE runbooks and incident management tooling.
- Security and governance components such as Vault Keycloak network policies RBAC admission controls image scanning SBOM tooling and audit logging.
- Infrastructure awareness covering GPU-backed nodes CPU-only fallback storage performance network isolation proxy patterns on-premise environments and dedicated landing zones.
Qualifications :
- 5 years in SRE platform engineering DevOps cloud infrastructure or operations roles with strong Kubernetes and Linux expertise.
- Proven experience building and operating production-grade engineering platforms with GitOps Infrastructure as Code observability and operational runbooks.
- Hands-on skills in Terraform Ansible Helm Python or shell scripting CI/CD runners private registries and secure configuration management.
- Good understanding of networking storage secrets access control monitoring backup disaster recovery and operational hardening in high-security environments.
- Comfortable supporting AI-enabled engineering workloads in sovereignty-driven contexts where isolation controlled data handling reliability and auditability are mandatory.
Additional Information :
What do we offer you
Work environment & flexibility
- International dynamic and collaborative environment.
- T-Social: social initiatives (sports community health ...).
- Hybrid work model (remote/on-site).
- Flexible working hours.
Growth & development
- Customized training: access to Coursera to learn whatever you want whenever you want.
- Weekly language classes (English & German).
- International Mentoring Sessions & Experience Days.
Compensation & benefits
- Flexible compensation plan (health insurance meal vouchers childcare transport).
- Telemedicine.
- Life and accident insurance.
- Social fund.
Wellbeing & time off
- 26 working days of vacation per year.
- Free access to specialist services (medical legal wellness).
- 100% salary coverage during medical leave.
And many more advantages of being part of T-Systems!
If you are looking for a new challenge do not hesitate to send us your CV! Please send CV in English. Join our team!
T-Systems Iberia will only process the CVs of candidates who meet the requirements specified for each offer.
Remote Work :
No
Employment Type :
Full-time
About Company
En T-Systems, encontrarás proyectos rompedores que suman al bienestar social y ecológico. Queremos dar la bienvenida a nuevos talentos como tú, que aporten ideas frescas, puntos de vista distintos, que acepten retos y un continuo aprendizaje, para crecer e impactar a la sociedad… ¡Tod ... View more