Dice logo
Dice
Posted 35 days agoVerified live 10h ago

Site Reliability Engineer, USG

Brief overview

Remote
$125k–$150k/yrStated range
4+ yrsMinimum
Clearance requiredU.S. government
Observability PlatformsPrometheusGrafanaOpenTelemetryGoogle Cloud PlatformKubernetesInfrastructure as CodeGitOpsPythonHigh-Availability Distributed Systems

About the company

Dice is the go-to career marketplace for tech professionals.

Job description

Summary

Aalyria is a technology company developing laser communications technology and temporospatial software-defined networking platforms for the aerospace industry. The Site Reliability Engineer will build and manage a production-grade observability platform for satellite, ground station, and distributed network systems, while defining reliability practices, automating infrastructure, and leading monitoring and incident response.

Responsibilities

  • Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry)
  • Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready
  • Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries)
  • Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g., ArgoCD)
  • Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying Google Cloud Platform and AWS environments
  • Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems

Skills

  • Active Top Secret (TS/SCI) security clearance
  • 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems
  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems
  • Strong production-level experience with Google Cloud Platform (Google Cloud Platform) and Kubernetes
  • Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD)
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems
  • U.S. citizen or national
  • U.S. lawful permanent resident ()
  • Refugee under 8 U.S.C. 1157
  • Asylee under 8 U.S.C. 1158
  • Be eligible to access export-controlled information without requiring an export authorization
  • Be eligible and reasonably likely to obtain the necessary export authorization from the appropriate U.S. government agency
  • Experience operating a multi-cloud environment, specifically Google Cloud Platform and AWS
  • Hands-on experience with GitLab CI for CI/CD pipelines
  • Working knowledge of service mesh technologies such as Istio or Linkerd
  • Familiarity with instrumenting applications written in Go and C++
  • An active Secret clearance, or higher, is preferred for this position
  • Experience with JVM observability (tuning, monitoring) for Java-based applications

Qualifications

Must Haves

  • Active Top Secret (TS/SCI) security clearance
  • 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems
  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems
  • Strong production-level experience with Google Cloud Platform (Google Cloud Platform) and Kubernetes
  • Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD)
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems
  • U.S. citizen or national
  • U.S. lawful permanent resident ()
  • Refugee under 8 U.S.C. 1157
  • Asylee under 8 U.S.C. 1158
  • Be eligible to access export-controlled information without requiring an export authorization
  • Be eligible and reasonably likely to obtain the necessary export authorization from the appropriate U.S. government agency

Nice to Haves

  • Experience operating a multi-cloud environment, specifically Google Cloud Platform and AWS
  • Hands-on experience with GitLab CI for CI/CD pipelines
  • Working knowledge of service mesh technologies such as Istio or Linkerd
  • Familiarity with instrumenting applications written in Go and C++
  • An active Secret clearance, or higher, is preferred for this position
  • Experience with JVM observability (tuning, monitoring) for Java-based applications

Benefits

  • Opportunities for professional development and advancement
  • Flexible working arrangements including hybrid remote/in-office schedules
  • 401(k)
  • Dental insurance
  • Vision insurance
  • Health insurance
  • Life insurance
  • Paid time off
  • Equity options
  • Innovative environment at a cutting-edge company shaping the future of aerospace communications
  • Collaborative, supportive, and inclusive workplace

More jobs like this