Summary
Aalyria is a technology company developing laser communications technology and temporospatial software-defined networking platforms for the aerospace industry. The Site Reliability Engineer will build and manage a production-grade observability platform for satellite, ground station, and distributed network systems, while defining reliability practices, automating infrastructure, and leading monitoring and incident response.
Responsibilities
- Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry)
- Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready
- Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries)
- Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g., ArgoCD)
- Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying Google Cloud Platform and AWS environments
- Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems
Skills
- Active Top Secret (TS/SCI) security clearance
- 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems
- Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems
- Strong production-level experience with Google Cloud Platform (Google Cloud Platform) and Kubernetes
- Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD)
- Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling
- Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems
- U.S. citizen or national
- U.S. lawful permanent resident ()
- Refugee under 8 U.S.C. 1157
- Asylee under 8 U.S.C. 1158
- Be eligible to access export-controlled information without requiring an export authorization
- Be eligible and reasonably likely to obtain the necessary export authorization from the appropriate U.S. government agency
- Experience operating a multi-cloud environment, specifically Google Cloud Platform and AWS
- Hands-on experience with GitLab CI for CI/CD pipelines
- Working knowledge of service mesh technologies such as Istio or Linkerd
- Familiarity with instrumenting applications written in Go and C++
- An active Secret clearance, or higher, is preferred for this position
- Experience with JVM observability (tuning, monitoring) for Java-based applications
Qualifications
Must Haves
- Active Top Secret (TS/SCI) security clearance
- 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems
- Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems
- Strong production-level experience with Google Cloud Platform (Google Cloud Platform) and Kubernetes
- Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD)
- Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling
- Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems
- U.S. citizen or national
- U.S. lawful permanent resident ()
- Refugee under 8 U.S.C. 1157
- Asylee under 8 U.S.C. 1158
- Be eligible to access export-controlled information without requiring an export authorization
- Be eligible and reasonably likely to obtain the necessary export authorization from the appropriate U.S. government agency
Nice to Haves
- Experience operating a multi-cloud environment, specifically Google Cloud Platform and AWS
- Hands-on experience with GitLab CI for CI/CD pipelines
- Working knowledge of service mesh technologies such as Istio or Linkerd
- Familiarity with instrumenting applications written in Go and C++
- An active Secret clearance, or higher, is preferred for this position
- Experience with JVM observability (tuning, monitoring) for Java-based applications
Benefits
- Opportunities for professional development and advancement
- Flexible working arrangements including hybrid remote/in-office schedules
- 401(k)
- Dental insurance
- Vision insurance
- Health insurance
- Life insurance
- Paid time off
- Equity options
- Innovative environment at a cutting-edge company shaping the future of aerospace communications
- Collaborative, supportive, and inclusive workplace