T3 Operations & Support Specialist — Network & Security (PID9067)
Job Summary
This is a remote position.
T3 Operations & Support Specialist Network & Security (PID9067)
- Contract / Freelance
- Full-time
- Remote with travel readiness required (Germany)
- Start: ASAP
About the role
We are working with a long-standing anchor client to source a T3 Operations & Support Specialist (Network & Security) for a large-scale cloud-native platform programme supporting a major energy transmission operator in Germany. The platform is a service-oriented hybrid cloud environment providing application teams with self-service capabilities to develop run and operate software products across private and public cloud infrastructure.
In this role you will provide Tier-3 operational ownership for Network & Security services within Local Production (DE) resolving complex connectivity and security incidents driving RCA and remediation across the network/security stack and ensuring operational readiness for all network and security changes.
What youll be doing
- Providing T3 operational ownership for Network & Security services: resolving complex incidents driving RCA and remediation across the full network/security stack
- Ensuring operational readiness for network/security changes: monitoring/alerting validation steps rollback strategies runbooks and maintenance procedures
- Supporting compliance-relevant operational controls (logging/monitoring evidence access enforcement patterns vulnerability handling coordination)
- Coordinating with platform and Kubernetes teams to resolve cluster and application impacts caused by network/security constraints
- Executing and improving standard operational procedures through automation to reduce toil and improve MTTR and stability
- Monitoring system health performance metrics and service availability across multi-tenant environments
- Identifying analysing and resolving incidents to minimise service disruption and triggering RCA and corrective actions
- Implementing monitoring and logging strategies to support audit and compliance requirements
- Performing routine security scans and remediating identified vulnerabilities
What youll need
- 5 years operating enterprise networks in private cloud or data centre production environments
- Proven experience implementing and leading Incident Problem Change and Release governance in production
- Proven incident response and troubleshooting skills across routing firewalling connectivity and service exposure patterns
- Strong understanding of security fundamentals and operational enforcement in production contexts
- Networking: WAN/LAN routers and firewalls (enterprise operations) from Cisco Juniper and Palo Alto
- Network and security software: ACI APSTRA and Panorama
- Connectivity services: tenant private networking and connectivity patterns supporting production platforms
- DNS and certificates: DNS operations and certificate lifecycle handling (issuance renewal rotation coordination) via Infoblox
- ITSM tooling: Jira Service Management Jira Confluence
- Fundamental understanding of core operations processes (Incident Change Problem management ITSM) and SRE concepts
- Experience gathering operational insights from monitoring/observability including SLI/SLA/SLO management and tracking
- Hands-on experience documenting procedures and enforcing clear runbooks and playbooks
- Hands-on experience with monitoring and logging tools (e.g. Prometheus Grafana Datadog Mimir Loki)
- Understanding of modern platform operations (Kubernetes/containers automation observability) sufficient to govern specialists
- Fluent English and German (C1 minimum in both)
Desirable
- Experience operating in regulated or high-availability industries (banking telco public sector healthcare)
- Experience with enterprise ITSM and change governance in regulated environments
- Experience with SRE practices (SLOs/SLIs error budgets) and reliability management
- Familiarity with enterprise DevOps toolchains (GitLab JFrog Artifactory Backstage Harness)
- IaC/GitOps: Terraform/OpenTofu Helm/ArgoCD