Summary
Eagle Eye Networks is the world’s largest AI cloud-native physical security company, and they are seeking a Site Reliability Engineer to enhance their operations. In this role, you will ensure high availability of their global platform while designing resilient systems and automating processes to minimize operational toil.
Responsibilities
- Build and maintain reliable, automated infrastructure across our cloud environments using Infrastructure as Code (IaC) tools
- Participate in the on-call rotation, assisting with troubleshooting, root-cause analysis, and follow-up actions to prevent recurring incidents
- Utilize strong scripting skills (Python, Bash, or Golang) to drive automation and reduce manual "toil"
- Apply best practices for monitoring and alerting using tools such as Prometheus/VictoriaMetrics and Grafana
- Work with cross-functional partners to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Support production readiness and contribute to the improvement of CI/CD tooling to empower our application teams
Skills
- 2+ years of experience as an SRE or in a related infrastructure-focused role
- Strong experience managing Linux systems in production environments
- Good working knowledge of Kubernetes or other container orchestration systems
- Solid abilities in Python or Bash; familiarity with Golang is a plus
- Proven ability to identify reliability issues and implement scalable improvements
- Hands-on experience with incident response and basic SLO/error-budget usage
- Familiarity with using or building internal tools that leverage LLMs to automate repetitive infrastructure tasks or incident response
Qualifications
Must Haves
- 2+ years of experience as an SRE or in a related infrastructure-focused role
- Strong experience managing Linux systems in production environments
- Good working knowledge of Kubernetes or other container orchestration systems
- Solid abilities in Python or Bash; familiarity with Golang is a plus
- Proven ability to identify reliability issues and implement scalable improvements
- Hands-on experience with incident response and basic SLO/error-budget usage
- Familiarity with using or building internal tools that leverage LLMs to automate repetitive infrastructure tasks or incident response
Benefits
- Competitive medical, vision, and dental plans (including company-offset premiums)
- 401k with company match
- Unlimited Paid Time Off (PTO) policy