Summary
NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits. They are seeking a Site Reliability Engineer II to ensure the availability, reliability, and performance of production platforms while collaborating closely with various engineering teams.
Responsibilities
- Serve as the first responder for production incidents by identifying, triaging, and resolving issues
- Monitor and respond to alerts generated by Datadog and other monitoring platforms
- Perform initial root cause analysis and escalate incidents according to defined SLAs
- Communicate incident status and resolution updates to internal stakeholders
- Partner with senior engineers to resolve complex production issues
- Continuously monitor application health, infrastructure performance, and system availability
- Configure and optimize monitoring dashboards and alert thresholds
- Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis
- Support containerized applications running in Kubernetes and Docker environments
- Participate in a weekday "Follow-the-Sun" production support model with global engineering teams
- Participate in an on-call rotation for critical production systems as needed
- Help maintain high availability and system uptime
- Develop automation scripts and operational tools using one or more of the following: Python, PowerShell, Bash, C#, Java
- Support CI/CD pipeline monitoring and deployment reliability
- Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks
- Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams
- Recommend improvements to monitoring, tooling, and operational processes
- Collaborate effectively with globally distributed engineering teams
- Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews
- Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST
Skills
- 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support
- Hands-on experience with production incident response, troubleshooting, and escalation
- Experience with Datadog or similar monitoring and observability platforms
- Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management
- Experience with Docker or other container technologies
- Working knowledge of SQL, MySQL, or NoSQL databases
- Ability to work effectively in high-volume, mission-critical production environments
- Strong analytical, troubleshooting, and problem-solving skills
- Excellent written and verbal communication skills
- Willingness to work weekday shifts as part of a global Follow-the-Sun support model
- Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP)
- Familiarity with CI/CD pipelines and deployment automation
- Experience with Helm Charts and Kubernetes deployments
- Knowledge of ITIL principles and Agile methodologies
- Experience supporting regulated environments such as healthcare or fintech
- Scripting or programming experience in Python, PowerShell, Bash, Java, or C#
Qualifications
Must Haves
- 3–5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support
- Hands-on experience with production incident response, troubleshooting, and escalation
- Experience with Datadog or similar monitoring and observability platforms
- Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management
- Experience with Docker or other container technologies
- Working knowledge of SQL, MySQL, or NoSQL databases
- Ability to work effectively in high-volume, mission-critical production environments
- Strong analytical, troubleshooting, and problem-solving skills
- Excellent written and verbal communication skills
- Willingness to work weekday shifts as part of a global Follow-the-Sun support model
Nice to Haves
- Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP)
- Familiarity with CI/CD pipelines and deployment automation
- Experience with Helm Charts and Kubernetes deployments
- Knowledge of ITIL principles and Agile methodologies
- Experience supporting regulated environments such as healthcare or fintech
- Scripting or programming experience in Python, PowerShell, Bash, Java, or C#
Benefits
- Work on technology that positively impacts millions of healthcare members.
- Join a collaborative, innovative, and supportive engineering culture.
- Exposure to modern cloud-native technologies and enterprise-scale infrastructure.
- Competitive compensation and comprehensive benefits.
- Unlimited Paid Time Off (PTO).
- Opportunities for career growth and professional development.
- Work with talented global engineering teams on challenging, high-impact projects.
- Maintain a healthy work-life balance while contributing to mission-critical platforms.
- Remote (US-Based Candidates Only)