Summary
Skydio is a leading U.S. drone company focused on autonomous flight technology. The Site Reliability Engineer will build, operate, troubleshoot, and scale production cloud infrastructure, with responsibility for Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability. The role also involves incident response, automation, and expanding infrastructure across regions and deployment environments.
Responsibilities
- Build, operate, and troubleshoot production Kubernetes/EKS clusters
- Perform Kubernetes upgrades, node rollouts, and cluster maintenance
- Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage
- Define and maintain infrastructure using Terraform
- Build and operate CI/CD and deployment infrastructure
- Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases
- Build monitoring, alerting, and observability for critical infrastructure
- Participate in on-call rotations and respond to production incidents
- Identify and solve infrastructure scaling and reliability problems
- Automate operational work using Python, Go, or similar languages
- Help expand infrastructure across new regions and deployment environments
Skills
- • 3+ years of experience as a Production Engineer, SRE, DevOps, SRE or equivalent infrastructure role
- • Strong hands-on experience operating Kubernetes, not simply deploying applications to existing clusters
- • Experience managing Kubernetes/EKS upgrades and production clusters
- • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases
- • Production experience with Terraform or similar infrastructure-as-code tooling
- • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins
- • Experience diagnosing production infrastructure and networking problems
- • Experience solving meaningful scaling or reliability challenges
- • Helm and GitOps experience
- • Datadog or similar observability tooling
- • PostgreSQL/database operations experience
- • Multi-region infrastructure experience
- • On-premises or disconnected deployment experience
- • Streaming or high-throughput distributed systems experience
Qualifications
Must Haves
- • 3+ years of experience as a Production Engineer, SRE, DevOps, SRE or equivalent infrastructure role
- • Strong hands-on experience operating Kubernetes, not simply deploying applications to existing clusters
- • Experience managing Kubernetes/EKS upgrades and production clusters
- • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases
- • Production experience with Terraform or similar infrastructure-as-code tooling
- • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins
- • Experience diagnosing production infrastructure and networking problems
- • Experience solving meaningful scaling or reliability challenges
Nice to Haves
- • Helm and GitOps experience
- • Datadog or similar observability tooling
- • PostgreSQL/database operations experience
- • Multi-region infrastructure experience
- • On-premises or disconnected deployment experience
- • Streaming or high-throughput distributed systems experience
Benefits
- Equity in the form of stock options
- Comprehensive benefits packages
- Relocation assistance may also be provided for eligible roles.
- Regular, full-time employees are eligible to enroll in the Company’s group health insurance plans.
- Paid vacation time
- Sick leave
- Holiday pay
- 401K savings plan