NationsBenefits logo
NationsBenefits
Posted 69 days agoVerified live 2d ago

Site Reliability Engineer II

Brief overview

Remote
3+ yrsMinimum
Site Reliability EngineeringDevOpsProduction Incident ResponseDatadogKubernetesDockerSQLMySQLNoSQL DatabasesPythonPowerShellBashC#JavaCI/CD PipelinesHelm ChartsMicrosoft Azure

About the company

NationsBenefits logo
NationsBenefitsNationsBenefits.com

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions.

Job description

Summary

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits. They are seeking a Site Reliability Engineer II to ensure the availability, reliability, and performance of production platforms while collaborating closely with various engineering teams.

Responsibilities

  • Serve as the first responder for production incidents by identifying, triaging, and resolving issues
  • Monitor and respond to alerts generated by Datadog and other monitoring platforms
  • Perform initial root cause analysis and escalate incidents according to defined SLAs
  • Communicate incident status and resolution updates to internal stakeholders
  • Partner with senior engineers to resolve complex production issues
  • Continuously monitor application health, infrastructure performance, and system availability
  • Configure and optimize monitoring dashboards and alert thresholds
  • Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis
  • Support containerized applications running in Kubernetes and Docker environments
  • Participate in a weekday "Follow-the-Sun" production support model with global engineering teams
  • Participate in an on-call rotation for critical production systems as needed
  • Help maintain high availability and system uptime
  • Develop automation scripts and operational tools using one or more of the following: Python, PowerShell, Bash, C#, Java
  • Support CI/CD pipeline monitoring and deployment reliability
  • Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks
  • Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams
  • Recommend improvements to monitoring, tooling, and operational processes
  • Collaborate effectively with globally distributed engineering teams
  • Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews
  • Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST

Skills

  • 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support
  • Hands-on experience with production incident response, troubleshooting, and escalation
  • Experience with Datadog or similar monitoring and observability platforms
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management
  • Experience with Docker or other container technologies
  • Working knowledge of SQL, MySQL, or NoSQL databases
  • Ability to work effectively in high-volume, mission-critical production environments
  • Strong analytical, troubleshooting, and problem-solving skills
  • Excellent written and verbal communication skills
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model
  • Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP)
  • Familiarity with CI/CD pipelines and deployment automation
  • Experience with Helm Charts and Kubernetes deployments
  • Knowledge of ITIL principles and Agile methodologies
  • Experience supporting regulated environments such as healthcare or fintech
  • Scripting or programming experience in Python, PowerShell, Bash, Java, or C#

Qualifications

Must Haves

  • 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support
  • Hands-on experience with production incident response, troubleshooting, and escalation
  • Experience with Datadog or similar monitoring and observability platforms
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management
  • Experience with Docker or other container technologies
  • Working knowledge of SQL, MySQL, or NoSQL databases
  • Ability to work effectively in high-volume, mission-critical production environments
  • Strong analytical, troubleshooting, and problem-solving skills
  • Excellent written and verbal communication skills
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model

Nice to Haves

  • Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP)
  • Familiarity with CI/CD pipelines and deployment automation
  • Experience with Helm Charts and Kubernetes deployments
  • Knowledge of ITIL principles and Agile methodologies
  • Experience supporting regulated environments such as healthcare or fintech
  • Scripting or programming experience in Python, PowerShell, Bash, Java, or C#

Benefits

  • Work on technology that positively impacts millions of healthcare members.
  • Join a collaborative, innovative, and supportive engineering culture.
  • Exposure to modern cloud-native technologies and enterprise-scale infrastructure.
  • Competitive compensation and comprehensive benefits.
  • Unlimited Paid Time Off (PTO).
  • Opportunities for career growth and professional development.
  • Work with talented global engineering teams on challenging, high-impact projects.
  • Maintain a healthy work-life balance while contributing to mission-critical platforms.
  • Remote (US-Based Candidates Only)

More jobs like this