Nscale logo
Nscale
Posted 18 days agoVerified live 9h ago

Operational Data & Observability Engineer

Brief overview

Remote
UndergradOr in progress
$145k–$180k/yrStated range
3+ yrsMinimum
Observability EngineeringPrometheusGrafanaDatadogNew RelicELK/Elastic StackSplunkPythonGoBashAWSAzureGoogle Cloud PlatformKubernetesDistributed TracingApplication Performance MonitoringTerraform

About the company

Nscale builds AI data centers and provides GPU cloud infrastructure that companies use to train, run, and scale large AI models.

Job description

Summary

Nscale is seeking an Operational Data & Observability Engineer to build and evolve monitoring, logging, and observability capabilities for production environments. The role focuses on developing observability solutions, managing operational data pipelines, supporting reliability and incident response, and administering observability platforms.

Responsibilities

  • Design and implement enterprise observability strategies across infrastructure, services, and applications
  • Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility
  • Build and maintain centralized logging and log analysis pipelines
  • Implement distributed tracing to improve visibility across microservices and complex application workflows
  • Establish performance baselines and develop anomaly detection strategies
  • Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems
  • Design and manage operational data pipelines that support monitoring and analytics
  • Develop APIs and integrations that enable operational data consumption across teams
  • Ensure data quality, consistency, retention, and cost-efficient storage practices
  • Troubleshoot production issues using monitoring, logging, and tracing data
  • Participate in an on-call rotation and support incident response activities
  • Create and maintain operational documentation, runbooks, and troubleshooting guides
  • Partner with engineering teams to improve platform reliability, scalability, and operational readiness
  • Continuously optimize observability infrastructure for performance and resilience
  • Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar technologies
  • Evaluate emerging observability tools and recommend improvements
  • Automate monitoring deployments, instrumentation, and platform configuration
  • Perform ongoing maintenance, upgrades, and lifecycle management of observability infrastructure
  • Improve platform visibility and operational health
  • Reduce Mean Time to Resolution (MTTR) during incidents
  • Increase alert quality while reducing unnecessary noise
  • Deliver highly available, scalable observability platforms
  • Improve engineering productivity through actionable monitoring and operational insights
  • Optimize observability infrastructure performance and cost efficiency
  • Participate in a rotating on-call schedule to support production environments
  • Support mission-critical systems with occasional after-hours or incident response responsibilities

Skills

  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering
  • Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent
  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions
  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent
  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM)
  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies
  • Solid understanding of application, infrastructure, networking, database, and storage performance monitoring
  • Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem-solving
  • Participate in a rotating on-call schedule to support production environments
  • Support mission-critical systems with occasional after-hours or incident response responsibilities
  • Experience supporting microservices-based architectures
  • Expertise across multiple observability platforms
  • Experience with incident management, root cause analysis, and post-incident reviews
  • Infrastructure as Code experience using Terraform, Ansible, or similar tools
  • Familiarity with eBPF or low-level Linux performance monitoring
  • Experience building custom telemetry, ETL, or operational data pipelines
  • Understanding of security monitoring, audit logging, and compliance requirements

Qualifications

Must Haves

  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering
  • Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent
  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions
  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent
  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM)
  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies
  • Solid understanding of application, infrastructure, networking, database, and storage performance monitoring
  • Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem-solving
  • Participate in a rotating on-call schedule to support production environments
  • Support mission-critical systems with occasional after-hours or incident response responsibilities

Nice to Haves

  • Experience supporting microservices-based architectures
  • Expertise across multiple observability platforms
  • Experience with incident management, root cause analysis, and post-incident reviews
  • Infrastructure as Code experience using Terraform, Ansible, or similar tools
  • Familiarity with eBPF or low-level Linux performance monitoring
  • Experience building custom telemetry, ETL, or operational data pipelines
  • Understanding of security monitoring, audit logging, and compliance requirements

Benefits

  • Hybrid or remote work arrangements available, depending on business needs.
  • This role may be eligible for bonus, equity, and/or commission programs.
  • Medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

More jobs like this