Enter a job title or keyword

Director, Data & Storage Reliability Engineering

ServiceNow


Job Location:

Santa Clara County, CA - USA

Monthly Salary: Not provided by the employer
Posted: 10 September 2026 (5 hours ago)
Application Deadline: 8 December 2026
Vacancies: 1 Vacancy

Job Summary

What you get to do in this role:

Team Management 

The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability resilience performance scalability observability and customer experience across ServiceNows database storage and supporting platform infrastructure. 

This leader will be responsible for building developing and scaling high-performing engineering teams focused on reliability engineering observability performance engineering diagnostics automation production analytics migration readiness resilience engineering and prevention engineering. 

Responsibilities include talent acquisition performance management career development succession planning objective setting coaching and prioritization of strategic initiatives. 

The role will establish a strong engineering-first culture centered on data-driven decision making continuous improvement operational excellence customer experience and systemic risk reduction. 

This position is accountable for identifying recurring failure patterns reliability risks performance bottlenecks scalability constraints migration challenges and operational inefficiencies across database services storage platforms cloud infrastructure and distributed application environments and driving engineering improvements that eliminate entire classes of issues before they impact customers. 

The Director will partner closely with Product Engineering Database Engineering Cloud Infrastructure Architecture Storage Engineering Support and Operations teams to ensure reliability observability performance and resilience considerations are incorporated throughout the software development lifecycle. 

The successful candidate will also partner closely with SWAT and Customer & Production Engineering teams to establish a continuous feedback loop between production operations and platform improvement. SWAT remains responsible for customer escalations production operations incident response and service restoration while this organization is responsible for identifying systemic opportunities defining engineering priorities and driving platform improvements that reduce future customer impact. 

The successful candidate will serve as the senior technical leader for complex reliability investigations customer-critical escalation reviews migration readiness assessments and platform improvement initiatives transforming production insights into long-term engineering outcomes. 

They will influence architectural decisions and technology investments by providing reliability expertise observability insights performance guidance and production-based evidence that improve platform resilience scalability efficiency and customer outcomes. 

This role requires a strong product mindset. The leader will treat reliability observability resilience performance and automation capabilities as products with roadmaps priorities adoption goals and measurable outcomes. They will be responsible for identifying the highest-value engineering opportunities prioritizing investments and driving adoption across multiple product and infrastructure organizations. 

Process and Procedures 

The successful candidate will establish scalable reliability engineering practices standards governance processes and operating models across the organization. 

They will drive adoption of observability standards reliability engineering frameworks resiliency assessments migration readiness practices diagnostics capabilities engineering guardrails and automation strategies. 

This leader will continuously evaluate incidents customer escalations migration outcomes platform telemetry performance trends capacity signals and operational data to identify systemic risks and drive long-term engineering improvements. 

The role will establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents recurring operational challenges migration learnings customer-impacting events and emerging platform risks. These insights will be used to prioritize engineering investments and platform improvements. 

The successful candidate will establish meaningful KPIs and engineering metrics that provide visibility into platform reliability resiliency performance operational efficiency customer experience engineering productivity and risk reduction. 

The successful candidate will leverage AI-powered tools analytics automation frameworks and production intelligence to identify emerging risks improve detection coverage accelerate engineering insights reduce operational toil and improve engineering productivity. 

They will use production telemetry incident learnings customer escalations migration outcomes observability data and operational trends to drive architectural improvements reliability investments platform standards and long-term engineering evolution. 

The Director will maintain a portfolio of reliability investments spanning observability performance diagnostics resilience automation and prevention balancing immediate customer needs with long-term platform strategy. 

The Director will champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis prevention resilience and continuous optimization. 

 


Qualifications :

To be successful in this role you have:

  • Experience in leveraging or critically thinking about how to integrate AI into work processes decision-making or problem-solving. This may include using AI-powered tools automating workflows analyzing AI-driven insights or exploring AIs potential impact on the function or industry.
  • Strong product mindset with demonstrated experience treating technical capabilities as products with roadmaps priorities customers adoption goals and measurable business outcomes. 

  • Experience translating production insights customer pain points operational challenges reliability risks and platform telemetry into prioritized engineering investments and long-term roadmaps. 

  • Experience partnering closely with production operations customer escalation teams reliability organizations and software engineering teams to drive systemic improvements based on operational learnings. 

  • Experience defining product strategies developing roadmaps prioritizing investments and aligning stakeholders across multiple organizations without direct authority. 

  • Experience operating a portfolio of engineering investments balancing short-term customer needs with long-term reliability performance scalability and resilience objectives. 

  • 15 years of experience in software engineering platform engineering reliability engineering infrastructure engineering database engineering distributed systems product management or large-scale SaaS environments. 

  • 8 years of engineering leadership experience including leading managers and globally distributed teams. 

  • Extensive experience leading Reliability Engineering Platform Engineering Database Engineering Infrastructure Engineering Production Engineering Performance Engineering or related technical organizations. 

  • Deep expertise in distributed systems databases storage technologies cloud infrastructure and large-scale SaaS architectures. 

  • Strong understanding of reliability engineering principles observability scalability resiliency operational excellence and performance engineering. 

  • Experience building and operating observability telemetry diagnostics reliability or performance capabilities at scale. 

  • Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements. 

  • Experience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomes. 

  • Experience driving engineering initiatives through data metrics customer impact analysis and measurable business outcomes. 

  • Experience leveraging AI technologies to improve decision-making analytics engineering workflows operational efficiency reliability insights automation or customer outcomes. 

  • Exceptional communication stakeholder management and leadership skills. 

  • Bachelors degree in Computer Science Engineering or a related technical field or equivalent practical experience. 

Desired Skills 

  • Previous Product Management experience in a platform infrastructure cloud database storage or SaaS environment. 

  • Experience applying product management disciplines such as roadmap planning prioritization customer-centric thinking outcome measurement and portfolio management to engineering organizations. 

  • Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads. 

  • Experience building and scaling Reliability Engineering Performance Engineering Platform Engineering SRE or Production Engineering organizations. 

  • Experience with observability platforms telemetry systems diagnostics frameworks and production analytics. 

  • Experience with migration readiness resiliency validation reliability testing operational risk reduction and large-scale cloud transformations. 

  • Experience leveraging AI technologies to improve anomaly detection forecasting incident analysis prioritization and engineering productivity. 

  • Strong understanding of distributed systems architecture cloud platform operations and hyperscale environments. 

  • Experience developing executive-facing reliability scorecards engineering metrics and business impact reporting. 

  • Experience influencing platform architecture database strategy storage strategy and long-term engineering roadmaps. 

  • Experience with Linux-based production environments and large-scale cloud infrastructure. 

  • Experience supporting enterprise database technologies such as MySQL MariaDB PostgreSQL Oracle SQL Server or cloud-native database platforms. 

  • Familiarity with ServiceNow platform architecture and large-scale SaaS operations. 

Why This Role Exists 

SWAT and Customer & Production Engineering teams are responsible for protecting customers when issues occur. Data & Storage Reliability Engineering exists to ensure fewer issues occur in the first place. 

This organization serves as the engineering and product partner to SWAT by transforming production insights customer escalations migration learnings telemetry performance data and operational experience into prioritized engineering investments.

 

 

JV20

For positions in this location we offer a base pay of $221200 - $387100 plus equity (when applicable) variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline and individual total compensation will vary based on factors such as qualifications skill level competencies and work location. We also offer health plans including flexible spending accounts a 401(k) Plan with company match ESPP matching donations a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.


Additional Information :

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color creed religion sex sexual orientation national origin or nationality ancestry age disability gender identity or expression marital status veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. 

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. 2026 Fortune Media IP Limited. All rights reserved. Used under license.


Remote Work :

No


Employment Type :

Full-time


About Company

Company Logo

Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. We’re growing fast, innovating even faster, and making an impact on our c ... View more

View Profile View Profile