Keybank logo
Keybank
Posted 10 days agoVerified live 2d ago

Resiliency Engineer

Brief overview

Remote
UndergradOr in progress
$63k–$96k/yrStated range
3+ yrsMinimum
89 H-1B approvalsDept. of Labor
29 green cardsCertified filings
Reliability EngineeringAutomation ScriptingPythonBashPowerShellTerraformAnsibleGoogle Cloud Platform (GCP)Microsoft AzureDisaster Recovery PlanningKubernetesCross-team Communication

About the company

Keybank logo
Keybankkey.com

A financial services company providing banking and advisory services.

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
89H-1B approved
96%approval rate
24new H-1B hires
29PERM certified
$121,450median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
202329
202428
202525
20267
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
202311
20245
20256
20266
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
20237
20249
202510
20263
Top sponsored roles
Quantitative Analytics Senior AssociateSenior Software EngineerQuantitative Analytics Lead AssociateQuantitative Analytics, Sr. AssociateQuantitative Analytics Senior Manager
Sponsored employees from
IndiaChinaTaiwan

Job description

Summary

KeyBank is seeking a Resiliency Engineer to strengthen the reliability, availability, and recoverability of its technology platforms across on-premises, hybrid, and cloud environments. The role develops automation, designs resiliency solutions, conducts recovery and chaos testing, and partners with technical and business teams to meet recovery objectives and regulatory requirements.

Responsibilities

  • Design and code automation that reduces operational toil and replaces manual, error-prone runbooks with orchestrated, auditable failover and recovery workflows
  • Develop and maintain infrastructure-as-code, scripts, and pipelines (e.g., Python, Bash, PowerShell, Terraform, Ansible) to provision, configure, and validate recovery environments
  • Build self-healing patterns, health checks, and automated validation that confirm recoverability before a disaster is ever declared
  • Partner with application and infrastructure teams to assess system architecture for reliability, redundancy, and recoverability against assigned system criticality
  • Define, measure, and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services
  • Ensure architecture is selected to meet Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets
  • Plan and execute disaster recovery tests and targeted fault-injection / chaos experiments (e.g., zonal failure, load-balancer, regional failover) to proactively expose weaknesses
  • Improve monitoring, alerting, and observability to detect service degradation early
  • Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs, aligning technical and business stakeholders toward clear outcomes
  • Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams
  • Produce clear, examiner-ready documentation and evidence, ensuring work aligns with KeyBank policies, standards, and regulatory requirements

Skills

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field—or equivalent work experience
  • Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations
  • Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and experience with infrastructure-as-code (e.g., Terraform, Ansible)
  • Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure)
  • Understanding of high-availability design: redundancy, replication, failover, and load balancing
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams
  • Experience with site reliability engineering principles and service-level management (SLIs, SLOs, error budgets)
  • Experience with disaster recovery planning, resiliency testing, or chaos / fault-injection engineering (e.g., Google FIT, Gremlin)
  • Familiarity with containers and orchestration (Kubernetes/GKE), CI/CD, and observability tooling (e.g., Dynatrace, Prometheus, Grafana, Splunk)
  • Experience with ServiceNow (ITOM / Business Continuity Management) or comparable orchestration platforms
  • Experience in a regulated industry or large, complex enterprise environment; familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301
  • Relevant certifications (e.g., cloud architect/engineer, Kubernetes, Linux, ITIL)

Qualifications

Must Haves

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field—or equivalent work experience
  • Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations
  • Hands-on coding and automation ability (e.g., Python, Bash, PowerShell) and experience with infrastructure-as-code (e.g., Terraform, Ansible)
  • Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure)
  • Understanding of high-availability design: redundancy, replication, failover, and load balancing
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams

Nice to Haves

  • Experience with site reliability engineering principles and service-level management (SLIs, SLOs, error budgets)
  • Experience with disaster recovery planning, resiliency testing, or chaos / fault-injection engineering (e.g., Google FIT, Gremlin)
  • Familiarity with containers and orchestration (Kubernetes/GKE), CI/CD, and observability tooling (e.g., Dynatrace, Prometheus, Grafana, Splunk)
  • Experience with ServiceNow (ITOM / Business Continuity Management) or comparable orchestration platforms
  • Experience in a regulated industry or large, complex enterprise environment; familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301
  • Relevant certifications (e.g., cloud architect/engineer, Kubernetes, Linux, ITIL)

Benefits

  • Eligibility for incentive compensation which may include production, commission, and/or discretionary incentives.
  • Flexible options when roles can be performed effectively in a mobile environment.

More jobs like this