Lead Infrastructure Engineer (Site Reliability Engineering) – Workplace Technology
Job Summary
About this role:
Wells Fargo is seeking a Lead Infrastructure Engineer to help drive reliability operational excellence automation and service resiliency across our Workplace Technology portfolio.
This role will provide technical leadership for the technologies that enable our employees to work collaborate and communicate every day. You will partner closely with engineering infrastructure security and architecture teams to improve service reliability reduce operational effort through automation strengthen observability and ensure a consistent end-user experience at enterprise scale.
Our Workplace Technology landscape includes:
- Microsoft 365 (Teams Exchange SharePoint OneDrive Copilot) Collaboration & Productivity Services
- Endpoint Management & Intune Software Distribution Services
- Entire PC Lifecycle Management
- Workplace Compute (Windows Endpoints VDI Windows 365)
- Compliance Security and Regulatory Services
This is a highly visible role that will influence the reliability strategy and operational maturity of Workplace Technology services globally.
In this role you will:
- Lead reliability and operational support for Workplace Technology platforms serving a global workforce.
- Drive the adoption of Site Reliability Engineering (SRE) practices including service health monitoring SLOs SLIs and operational metrics.
- Lead response and recovery efforts during major incidents coordinating across technology teams to restore services and minimize business impact.
- Drive root cause analysis and continuous improvement initiatives to improve service stability and prevent recurring issues.
- Champion automation self-healing capabilities and process simplification to reduce operational toil.
- Build and enhance observability capabilities using monitoring telemetry logging and user experience insights.
- Partner with engineering teams on platform modernization cloud adoption resiliency improvements and operational readiness.
- Establish operational governance risk controls and compliance practices across supported services.
- Lead capacity availability and performance management activities to ensure services scale effectively.
- Provide technical leadership mentoring and guidance to engineers and support teams operating in a global environment.
Required Qualifications:
- 5 years of experience in Systems Operations Infrastructure Engineering Site Reliability Engineering or Technology Operations.
- Experience supporting large-scale enterprise platforms with global operational responsibilities.
- Strong background in incident management problem management and service reliability.
- Proven track record of driving automation operational improvements and service optimization initiatives.
- Experience supporting critical production environments with high availability and resiliency requirements.
- Strong understanding of SRE principles operational excellence and modern support models.
Desired Qualifications:
Workplace Technology Experience
Hands-on experience with one or more of the following:
- Microsoft 365 (Exchange Online Teams SharePoint OneDrive Copilot)
- Windows Endpoints Windows 365 Intune and Endpoint Management
- Software Distribution and Endpoint Management platforms
Reliability & Observability
- Experience defining SLOs SLIs operational KPIs and reliability measures.
- Experience with enterprise monitoring and observability platforms such as Splunk ThousandEyes Azure Monitor Grafana Dynatrace Aternity or AppDynamics.
- Knowledge of service resiliency disaster recovery capacity planning and performance management.
Automation & Cloud
- Automation experience with PowerShell Python Ansible Terraform CI/CD REST APIs SQL Queries HTML Coding etc.
- Working knowledge of Microsoft Azure and hybrid cloud environments.
- Experience implementing automation and self-healing solutions at scale.
Job Expectations:
- Serve as a technical leader for Workplace Technology Operations with good knowledge on Infrastructure elements.
- Improve service reliability stability and end-user experience across Workplace Technology platforms.
- Drive observability automation and operational excellence initiatives.
- Partner with engineering teams to embed reliability and operational readiness into platform design.
- Support a 24x7 global environment through strong operational practices and effective incident management.
- Ensure adherence to enterprise security audit compliance and risk requirements.
- Foster a culture of ownership continuous improvement innovation and engineering excellence.
We are looking for a Lead Systems Operations Engineer to help drive reliability operational excellence automation service resiliency and continuous improvement across our Workplace Technology portfolio.
You will establish and promote SRE engineering standards and best practices across observability SLIs/SLOs error budgets incident and problem management automation and operational resilience. You will leverage telemetry and data-driven insights to proactively identify risks reduce incident frequency and recurrence and drive continuous reliability.
A key focus of your role will be improving the stability of workplace technology platforms while modernizing operational workflows through Agentic AI intelligent automation and low-code/no-code solutions. You will provide technical leadership for the development and adoption of AI-assisted incident triage and root-cause analysis automated workflows proactive remediation and self-healing capabilities designed to reduce operational toil and improve the end-user experience. You will partner closely with engineering infrastructure security and architecture teams to improve service reliability reduce operational effort through automation strengthen observability and ensure a consistent end-user experience at enterprise scale.
Posting End Date:
7 Sep 2026*Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation gender identity national origin disability status as a protected veteran or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit Market Financial Crimes Operational Regulatory Compliance) which includes effectively following and adhering to applicable Wells Fargo policies and procedures appropriately fulfilling risk and compliance obligations timely and effective escalation and remediation of issues and making sound risk decisions. There is emphasis on proactive monitoring governance risk identification and escalation as well as making sound risk decisions commensurate with the business units risk appetite and all risk and compliance program requirements.
Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates including women persons with disabilities aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.
Applicants with Disabilities
To request a medical accommodation during the application or interview process visitDisability Inclusion at Wells Fargo.
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.
Required Experience:
IC
About Company
Whether you’re just beginning your career or taking it to the next level, we have an opportunity for you.