Enter a job title or keyword

Resiliency Engineer

KeyBank


Job Location:

Albany, GA - USA

Yearly Salary: USD 63000 - 96000
Posted: 28 September 2026 (13 hours ago)
Application Deadline: 26 December 2026
Vacancies: 1 Vacancy

Job Summary

Location:

555 Patroon Creek Boulevard Albany New York

Position Summary

The Resiliency Engineer designs builds and continuously improves the reliability availability and recoverability of KeyBanks technology platforms across on-premises hybrid and cloud (GCP/Azure) environments. Applying software engineering discipline to operations this role engineers systems to meet defined recovery objectives automates recovery and validation and proves that critical services can withstand and recover from failure.
This is a hands-on engineering role for a builder who can write automation reason about distributed-system failure modes and facilitate across application infrastructure architecture cloud and line-of-business teams. The ideal candidate is equally comfortable in code in a design review and in a command center during a live recovery exercise.

Key Responsibilities

Automation & Engineering

  • Design and code automation that reduces operational toil and replaces manual error-prone runbooks with orchestrated auditable failover and recovery workflows.
  • Develop and maintain infrastructure-as-code scripts and pipelines (e.g. Python Bash PowerShell Terraform Ansible) to provision configure and validate recovery environments.
  • Build self-healing patterns health checks and automated validation that confirm recoverability before a disaster is ever declared.

Reliability & Resiliency Design

  • Partner with application and infrastructure teams to assess system architecture for reliability redundancy and recoverability against assigned system criticality.
  • Define measure and support Service Level Indicators (SLIs) Service Level Objectives (SLOs) and error budgets for critical services.
  • Ensure architecture is selected to meet Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.

Testing Chaos & Validation

  • Plan and execute disaster recovery tests and targeted fault-injection / chaos experiments (e.g. zonal failure load-balancer regional failover) to proactively expose weaknesses.
  • Improve monitoring alerting and observability to detect service degradation early.

Facilitation & Partnership

  • Facilitate resiliency and architecture reviews tabletop exercises and cross-team recovery walkthroughs aligning technical and business stakeholders toward clear outcomes.
  • Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams.
  • Produce clear examiner-ready documentation and evidence ensuring work aligns with KeyBank policies standards and regulatory requirements.
Required Qualifications
  • Bachelors degree in Computer Science Information Technology Engineering or related fieldor equivalent work experience.
  • Demonstrated experience in reliability engineering DevOps infrastructure or technology operations.
  • Hands-on coding and automation ability (e.g. Python Bash PowerShell) and experience with infrastructure-as-code (e.g. Terraform Ansible).
  • Working knowledge of BOTH on-premises infrastructure (compute storage network virtualization databases) AND public cloud platforms (GCP and/or Azure).
  • Understanding of high-availability design: redundancy replication failover and load balancing.
  • Strong facilitation analytical problem-solving and written/verbal communication skills with the ability to influence across teams.
Preferred Qualifications
  • Experience with site reliability engineering principles and service-level management (SLIs SLOs error budgets).
  • Experience with disaster recovery planning resiliency testing or chaos / fault-injection engineering (e.g. Google FIT Gremlin).
  • Familiarity with containers and orchestration (Kubernetes/GKE) CI/CD and observability tooling (e.g. Dynatrace Prometheus Grafana Splunk).
  • Experience with ServiceNow (ITOM / Business Continuity Management) or comparable orchestration platforms.
  • Experience in a regulated industry or large complex enterprise environment; familiarity with FFIEC NIST SP 800-34/CSF or ISO 22301.
  • Relevant certifications (e.g. cloud architect/engineer Kubernetes Linux ITIL).

COMPENSATION AND BENEFITS

This position is eligible to earn a base salary in the range of $63000.00 - $96000.00 annually. Placement within the pay range may differ based upon various factors including but not limited to skills experience and geographic location. Compensation for this role also includes eligibility for incentive compensation which may include production commission and/or discretionary incentives.

Please click here for a list of benefits for which this position is eligible.

Key has implemented an approach to employee workspaces which prioritizes in-office presence while providing flexible options in circumstances where roles can be performed effectively in a mobile environment.

Job Posting Expiration Date: 10/30/2026 KeyCorp is an Equal Opportunity Employer committed to sustaining an inclusive culture. All qualified applicants will receive consideration for employment without regard to race color religion sex sexual orientation gender identity national origin age genetic information pregnancy disability veteran status or any other characteristic protected by law.

Qualified individuals with disabilities or disabled veterans who are unable or limited in their ability to apply on this site may request reasonable accommodations by emailing

#LI-Remote

Required Experience:

IC