Summary
KeyBank is seeking a Resiliency Engineer to strengthen the reliability, availability, and recoverability of its technology platforms across on-premises, hybrid, and cloud environments. The role develops automation, designs resiliency solutions, conducts recovery and chaos testing, and partners with technical and business teams to meet recovery objectives and regulatory requirements.
Responsibilities
- Design and code automation that reduces operational toil and replaces manual, error-prone runbooks with orchestrated, auditable failover and recovery workflows
- Develop and maintain infrastructure-as-code, scripts, and pipelines (e.g., Python, Bash, PowerShell, Terraform, Ansible) to provision, configure, and validate recovery environments
- Build self-healing patterns, health checks, and automated validation that confirm recoverability before a disaster is ever declared
- Partner with application and infrastructure teams to assess system architecture for reliability, redundancy, and recoverability against assigned system criticality
- Define, measure, and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services
- Ensure architecture is selected to meet Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets
- Plan and execute disaster recovery tests and targeted fault-injection / chaos experiments (e.g., zonal failure, load-balancer, regional failover) to proactively expose weaknesses
- Improve monitoring, alerting, and observability to detect service degradation early
- Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs, aligning technical and business stakeholders toward clear outcomes
- Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams
- Produce clear, examiner-ready documentation and evidence, ensuring work aligns with KeyBank policies, standards, and regulatory requirements
Skills
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field—or equivalent work experience
- Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations
- Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and experience with infrastructure-as-code (e.g., Terraform, Ansible)
- Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure)
- Understanding of high-availability design: redundancy, replication, failover, and load balancing
- Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams
- Experience with site reliability engineering principles and service-level management (SLIs, SLOs, error budgets)
- Experience with disaster recovery planning, resiliency testing, or chaos / fault-injection engineering (e.g., Google FIT, Gremlin)
- Familiarity with containers and orchestration (Kubernetes/GKE), CI/CD, and observability tooling (e.g., Dynatrace, Prometheus, Grafana, Splunk)
- Experience with ServiceNow (ITOM / Business Continuity Management) or comparable orchestration platforms
- Experience in a regulated industry or large, complex enterprise environment; familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301
- Relevant certifications (e.g., cloud architect/engineer, Kubernetes, Linux, ITIL)
Qualifications
Must Haves
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field—or equivalent work experience
- Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations
- Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and experience with infrastructure-as-code (e.g., Terraform, Ansible)
- Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure)
- Understanding of high-availability design: redundancy, replication, failover, and load balancing
- Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams
Nice to Haves
- Experience with site reliability engineering principles and service-level management (SLIs, SLOs, error budgets)
- Experience with disaster recovery planning, resiliency testing, or chaos / fault-injection engineering (e.g., Google FIT, Gremlin)
- Familiarity with containers and orchestration (Kubernetes/GKE), CI/CD, and observability tooling (e.g., Dynatrace, Prometheus, Grafana, Splunk)
- Experience with ServiceNow (ITOM / Business Continuity Management) or comparable orchestration platforms
- Experience in a regulated industry or large, complex enterprise environment; familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301
- Relevant certifications (e.g., cloud architect/engineer, Kubernetes, Linux, ITIL)
Benefits
- Eligibility for incentive compensation which may include production, commission, and/or discretionary incentives.
- Flexible options when roles can be performed effectively in a mobile environment.