Enter a job title or keyword

Site Reliability Engr II

Honeywell


Job Location:

Bengaluru - India

Monthly Salary: Not provided by the employer
Posted: 8 August 2026 (17 days ago)
Application Deadline: 5 November 2026
Vacancies: 1 Vacancy

Job Summary

Description

Job Title:Senior Site Reliability Engineer

Location: Bangalore

We are seeking a highly technical SRE Engineer to design build and maintain fault-tolerant scalable and highly available distributed systems. You will champion SRE best practices reduce manual operations (toil) via automation and partner with product development squads to embed reliability into the software delivery lifecycle

Your role will include overseeing supervising and reviewing tasks performed by team members to ensure effective execution of work; managing end-to-end processes and projects for both internal and external clients with responsibility for timely and accurate delivery; issuing clear instructions and directions to team members on tasks to be performed; and mentoring and guiding junior colleagues to Support their skill development professional growth and overall success.



Responsibilities

Key Responsibilities:

Reliability & Availability

  • Ensure high availability and uptime of production services.
  • Define and manage Service Level Objectives (SLOs) Service Level Indicators (SLIs) and Service Level Agreements (SLAs).
  • Design and implement disaster recovery (DR) strategies including RTO/RPO targets.
  • Conduct capacity planning and scalability assessments.

Monitoring & Observability

  • Implement and maintain monitoring logging and alerting systems.
  • Create dashboards to track system health and performance.
  • Improve observability using tools such as Prometheus Grafana Azure Monitor Dynatrace or Elastic.
  • Proactively detect investigate and resolve system issues.
  • Incident Management
  • Participate in on-call rotations and incident response.
  • Lead troubleshooting during service disruptions.
  • Perform Root Cause Analysis (RCA) and drive corrective actions.
  • Reduce Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR).

Automation & Engineering

  • Develop automation to eliminate repetitive operational tasks.
  • Build self-healing and auto-scaling capabilities.
  • Create and maintain Infrastructure as Code (IaC).
  • Improve CI/CD pipelines and deployment reliability.

Cloud & Infrastructure Management

  • Manage cloud environments and production infrastructure.
  • Support Kubernetes/AKS clusters databases networking storage and messaging services.
  • Optimize infrastructure costs and resource utilization.
  • Ensure security compliance and operational best practices.

Collaboration

  • Partner with software development teams throughout the software lifecycle.
  • Participate in architecture reviews and non-functional requirement (NFR) assessments.
  • Review reliability risks and recommend improvements.
  • Support release planning and production readiness reviews.

Required Technical Skills

Infrastructure & Cloud

  • Azure/AWS or GCP
  • Kubernetes / AKS / OpenShift
  • Linux administration
  • Networking fundamentals (TCP/IP DNS Load Balancing)
  • Database administration basics (SQL/NoSQL)
  • Automation & DevOps
    • Terraform ARM Bicep or CloudFormation
  • CI/CD tools (Azure DevOps GitHub Actions Jenkins GitLab)
  • Infrastructure as Code (IaC)
  • Configuration management tools such as Ansible
  • Programming such as Python Jave Sheel scripting Go

Monitoring & Observability

  • Grafana
  • Prometheus
  • Dynatrace
  • Elastic Stack
  • Azure Monitor / Log Analytics


Qualifications

Experience Level: 5 yrs

Required Qualifications

  • Education: Bachelors/Masters degree in Computer Science Information Technology or equivalent practical experience.
  • Experience supporting large-scale cloud-native applications.
  • Experience with incident management and production support.
  • Understanding of SRE concepts such as:
    • Error Budgets
    • SLI/SLO/SLA
    • Chaos Engineering
    • High Availability
    • Reliability Engineering

Success Metrics

  • Service Availability (% Uptime)
  • SLA/SLO Compliance
  • MTTR / MTTD
  • Deployment Success Rate
  • Production Incident Reduction
  • Operational Cost Optimization
  • Automation Coverage
  • Customer Experience Metrics




About Company

Company Logo

Honeywell helps organizations solve the world's most complex challenges in automation, the future of aviation and energy transition. As a trusted partner, we provide actionable solutions and innovation through our Aerospace Technologies, Building Automation, Energy and Sustainability ... View more

View Profile View Profile