Enter a job title or keyword

Senior Manager, Data & Storage Reliability Engineering

ServiceNow


Job Location:

Dublin - Ireland

Monthly Salary: Not provided by the employer
Posted: 23 September 2026 (12 hours ago)
Application Deadline: 21 December 2026
Vacancies: 1 Vacancy

Job Summary

ServiceNow is seeking an experienced Sr Manager for our Data & Storage Reliability Engineering team. 

This leader will drive engineering excellence across prevention engineering reliability engineering observability incident learnings diagnostics automation capacity planning and platform risk reduction. The role requires deep technical expertise in database and distributed systems architectures large-scale SaaS production environments and customer-facing operations combined with strong people leadership and execution rigor. 

The ideal candidate has experience leading high-performing engineering teams responsible for identifying recurring production patterns converting incident and escalation learnings into durable engineering improvements strengthening observability and building sustainable solutions that improve platform resilience at scale. 

You should have experience with large-scale web applications database platforms distributed systems Linux-based production environments and a strong problem-solving mindset for reliability automation diagnostics observability and prevention. Qualified candidates will be responsible for leading a team of highly skilled engineers that push the limits of scalability resiliency and operational excellence. 

Do you 

  • Have experience leading teams of engineers and developing people 
  • Enjoy problem solving and using an analytical mindset to understand why systems fail and how to prevent repeat issues 
  • Have a technical background in roles including database engineering reliability engineering systems/cloud engineering SRE DevOps or production engineering 
  • Know Linux operating systems databases observability diagnostics and production troubleshooting well enough to guide engineers through complex investigations 
  • Have an attitude of continuous improvement and a passion for removing inefficient repetitive or reactive processes through automation and engineering prevention 

Answer yes to these questions and we want to hear from you. Hit the Apply button and lets have a chat about the role and your skills and experiences. 

Lets start with the role 

As a Sr Manager of the Data & Storage Reliability Engineering team your responsibilities will be: 

  • Define and execute the team-level strategy for prevention engineering reliability observability resilience and operational risk reduction across large-scale production environments. 
  • Lead initiatives that turn production signals incident learnings customer escalations migration outcomes and platform telemetry into durable engineering improvements. 
  • Partner closely with SWAT and Customer & Production Engineering to establish a continuous feedback loop between production operations and platform improvement. 
  • Drive improvements in observability diagnostics automation reliability reviews resiliency validation migration readiness and engineering guardrails. 
  • Identify recurring failure patterns reliability risks observability gaps operational inefficiencies scalability constraints and performance bottlenecks and drive action to reduce future customer impact. 
  • Establish reliability resilience observability automation and prevention goals for critical database and storage services. 
  • Champion proactive monitoring production analytics and automation to improve operational health and reduce repetitive manual work. 
  • Lead deep root cause analysis and ensure sustainable corrective actions are implemented for recurring issues and customer-impacting events. 
  • Partner with engineering leaders to influence database storage reliability observability and platform architecture priorities based on production evidence. 
  • Build and develop a world-class team of reliability prevention observability and platform engineers. 
  • Own team management recruitment career development objective setting project prioritization onboarding and performance reviews. 
  • Manage an engineering team that supports production-facing work including on-call or escalation participation where required. 
  • Drive a culture of intolerance for repetitive manual activities by promoting automation self-service diagnostics guardrails and scalable engineering practices. 
  • Drive initiatives with partner teams to improve the reliability resilience scalability and operational efficiency of the ServiceNow application and platform. 
  • Act as part of the escalation and crisis management ecosystem by helping convert immediate recovery learnings into sustainable engineering prevention. 
  • Analyze and evaluate existing processes to drive continuous improvement operational efficiency and prevention-oriented engineering practices. 
  • Provide training documentation dashboards playbooks and support to partner teams that interface with the Data & Storage Reliability Engineering team. 
  • Onboard new hires new technologies new systems and new automations into the team to enable successful execution and scale. 

Qualifications :

  • Experience in leveraging or critically thinking about how to integrate AI into work processes decision-making or problem-solving using AI-powered tools automating workflows analyzing AI-driven insights or exploring AIs potential impact on the function. 

  • 10 years of experience in database engineering reliability engineering distributed systems platform engineering infrastructure engineering production engineering or large-scale SaaS platform operations. 

  • 4 years of engineering leadership experience including leading engineers and cross-functional or distributed teams. 

  • Experience leading Reliability Engineering Database Engineering Platform Engineering Infrastructure Engineering Production Engineering Performance Engineering or related technical teams. 

  • Strong expertise in database technologies operating system performance distributed systems cloud-native architectures and large-scale production environments. 

  • Solid understanding of reliability engineering observability diagnostics root cause analysis capacity planning scalability engineering resiliency automation and operational excellence. 

  • Experience translating production insights customer escalations incident learnings platform telemetry and recurring operational challenges into prioritized engineering work. 

  • Experience designing and improving observability diagnostics reliability reviews migration readiness checks resiliency validation automation and engineering guardrails. 

  • Strong understanding of tuning and troubleshooting across database operating system storage network and application layers. 

  • Proven experience identifying and resolving complex reliability scalability performance efficiency and operational bottlenecks in large-scale distributed environments. 

  • Experience leading customer-critical investigations involving reliability capacity performance scalability resilience or operational risk challenges. 

  • Experience leveraging observability and telemetry platforms to analyze system behavior and drive platform improvements. 

  • Experience partnering with software engineering infrastructure production operations and escalation organizations to improve platform reliability scalability efficiency and production readiness. 

  • Experience driving engineering initiatives through data metrics incident learnings telemetry benchmarking and measurable outcomes. 

  • Strong communication stakeholder management and leadership skills. 

  • Bachelors degree in Computer Science Engineering or a related technical field or equivalent practical experience. 

Desired Skills 

  • Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads. 

  • Experience growing Reliability Engineering Database Engineering Platform Engineering Performance Engineering Scalability Engineering or Production Engineering teams. 

  • Experience with observability platforms telemetry systems diagnostics frameworks production analytics reliability scorecards engineering metrics and impact reporting. 

  • Experience with reliability reviews resiliency validation migration readiness workload simulation capacity forecasting prevention programs and operational risk reduction frameworks. 

  • Experience leveraging AI technologies to improve anomaly detection forecasting incident analysis prioritization operational efficiency and engineering productivity. 

  • Understanding of distributed systems architecture cloud platform operations Linux-based production environments and hyperscale environments. 

  • Experience contributing to platform architecture database strategy reliability investments scalability roadmaps and long-term engineering improvements. 

  • Experience with performance testing benchmarking workload simulation and capacity modeling as part of broader reliability and prevention engineering programs. 

  • Experience supporting enterprise database technologies such as MySQL MariaDB PostgreSQL Oracle SQL Server or cloud-native database platforms. 

  • Familiarity with ServiceNow platform architecture and large-scale SaaS operations. 


Additional Information :

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation national origin age disability gender identity  veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. 

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. 2026 Fortune Media IP Limited. All rights reserved. Used under license.


Remote Work :

No


Employment Type :

Full-time


About Company

Company Logo

Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. We’re growing fast, innovating even faster, and making an impact on our c ... View more

View Profile View Profile