Enter a job title or keyword

Senior Site Reliability Engineer

Mirantis


Job Location:

Hyderabad - India

Monthly Salary: Not provided by the employer
Posted: 22 September 2026 (20 hours ago)
Application Deadline: 20 December 2026
Vacancies: 1 Vacancy

Department:

Engineering

Job Summary

We are looking for a senior Kubernetes-focused DevOps/SRE engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments pipelines and infrastructure tooling that our engineering teams across the US Europe and APAC  depend on to ship daily.

This role spans both sides of the line. You will make our development test and pre-production clusters fast and reproducible harden the Helm and CI/CD path from commit to release and carry operational ownership including on-call for one of our smaller customer-facing production regions under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated the person other teams consult including the teams running our larger regions. Working within an agile framework you will directly shape how quickly and safely changes reach production and be accountable for how they behave once there.

Main Responsibilities:

  • Own the Kubernetes footprint across development CI pre-production and one customer-facing production region local kind clusters shared dev and QA environments and multi-cluster/multi-region topologies.

  • Operate your production region against defined SLOs: capacity and upgrade planning patching backup and restore disaster recovery drills and participation in an on-call rotation.

  • Lead incident response for your region detection mitigation customer-impact assessment root-cause analysis and blameless postmortems that feed fixes back into the platform.

  • Build and maintain Helm charts and umbrella releases for the platforms services and dependencies including versioning values hygiene and upgrade paths.

  • Own the CI/CD pipelines end to end build test image publishing chart packaging release cutting and hotfix/backport flows.

  • Automate environment bootstrap and seeding so any engineer can bring up a full stack control plane identity gateway database workflow engine with one command.

  • Operate and troubleshoot the supporting stack across test and production: PostgreSQL Temporal Keycloak API gateway message broker and observability components.

  • Build observability and diagnostics metrics dashboards alerting log and audit access that serve both engineering environments and production operations.

  • Consult with product teams and with the teams operating our larger regions on deployment topology GPU and resource scheduling RBAC networking and failure modes; validate upgrade and migration procedures and hand over runbooks.

  • Enforce security and tenant isolation in production: least-privilege access secret handling certificate and TLS lifecycle image and dependency scanning and audit evidence for compliance reviews.

  • Drive infrastructure as code and repeatability no snowflake environments no undocumented manual steps.

  • Mentor engineers on Kubernetes and operational practice and raise the teams bar through review and documentation.


Qualifications :

Required Skills/Abilities: 

  • 10 years in DevOps SRE platform or infrastructure engineering including production ownership of customer-facing Kubernetes environments.

  • Expert-level Kubernetes: workloads networking storage RBAC resource management CRDs and operators and cluster upgrades able to debug from kubectl and cluster internals rather than dashboards alone.

  • Proven incident response under SLA pressure on-call rotations escalation paths postmortems and follow-through on corrective action.

  • Strong CI/CD engineering pipelines as code reproducible builds artifact and release management (GitHub Actions or equivalent).

  • Solid scripting and automation ability and enough Go familiarity to read service code trace a failure into it and file a precise bug.

  • Experience running the stateful supporting stack relational databases identity providers gateways and message brokers in Kubernetes including backup restore and upgrade.

  • Track record as a technical consultant to other engineering teams: clear runbooks design feedback and incident write-ups across global time zones (strong written English).

Must Have

We expect deep production experience in several of these and the engineering fundamentals to learn the rest rapidly.

  • Kubernetes Native: Kubernetes at scale Cluster API controllers/operators Docker and Helm chart authoring and lifecycle management.

  • Delivery: GitHub Actions or equivalent CI/CD container registries versioned release and backport workflows.

  • Infrastructure as Code: Terraform Ansible or equivalent plus GitOps tooling (Argo CD Flux).

  • Identity & API Management: Keycloak and API gateway operation routing plugins TLS rate limiting.

  • Data & Messaging: PostgreSQL operations and migrations plus streaming/message-broker platforms (Kafka or equivalent).

  • Observability: Prometheus Grafana centralized logging and alerting tied to SLOs.

  • Cloud: AWS networking IAM load balancing and managed Kubernetes

Nice to Have

  • k0s or k0rdent ecosystem experience.

  • GPU infrastructure on Kubernetes device plugins node feature discovery scheduling and sharing of accelerators.

  • Temporal operations namespaces workers schema upgrades.

  • Bare-metal provisioning (Metal³ / BareMetalHost) or on-prem/OpenStack environments.

  • Multi-region topologies service mesh or cross-cluster networking.

  • Python for test harnesses and automation; experience with pytest-based E2E suites.

  • Load and performance testing of API platforms.

  • Policy enforcement (OPA/Kyverno) secret management and pen-test remediation.

  • OpenTelemetry distributed tracing or formal SLO/error-budget practice.

  • Compliance exposure SOC 2 ISO 27001 or similar audit support.

  • CNCF open-source contributions.

Education and Experience:

  • Bachelors degree in Computer Science & Engineering or related field or 10 years related experience.


Additional Information :

What does Mirantis offer you

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate talented and engaging colleagues helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge open-source innovation;
  • Thrive in the high-energy environment of a young company where openness collaboration risk-taking and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings happy hours hackathons and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!


Remote Work :

Yes


Employment Type :

Full-time


About Company

Mirantis is an open cloud company that helps organizations achieve digital self determination by giving them complete control over their strategic infrastructure. The company combines intelligent automation and cloud-native expertise for managing and operating virtual machines, contai ... View more

View Profile View Profile