Summary
SingleStore delivers a cloud-native distributed SQL database that powers data-intensive applications with real-time transactional and analytical capabilities. The Software Engineer will join the Observability Team to design and deliver scalable observability capabilities across distributed systems, cloud infrastructure, database technology, and AI-powered observability. The role includes building telemetry pipelines, alerting and visualization features, optimizing data storage and queries, and supporting reliable production operations.
Responsibilities
- Design and implement scalable observability features for traces, logs, and metrics — spanning ingestion, processing, storage, and visualization
- Work across control plane and data plane components in a multi-cloud environment (AWS, GCP, Azure), ensuring reliable operation and data consistency at scale
- Build high-throughput data pipelines that process telemetry data using OpenTelemetry Collector and related open-source tooling
- Develop and maintain alerting capabilities with Alertmanager, enabling customers to define, tune, and manage alerts with routing, inhibition, and notification management
- Optimize time-series data storage and query performance using SingleStore DB, handling high-cardinality data and complex analytical queries
- Contribute to data visualization dashboards in Grafana, creating intuitive customer-facing experiences to explore telemetry data
- Collaborate closely with Product Management to translate customer and business requirements into robust technical solutions
- Investigate and resolve difficult issues in production and development environments, debugging data synchronization across distributed systems and cloud providers
- Participate in on-call rotations to ensure system reliability and respond to incidents promptly
Skills
- 2+ years of professional software development experience building distributed systems or backend services
- Strong proficiency in Go (Golang) — experience with Rust, Python, or C++ is also valuable
- Deep understanding of distributed systems concepts: scalability, consistency, high availability, concurrency, and failure modes
- Familiarity with distributed systems managed via Kubernetes
- Demonstrated ability to design and build reliable, high-performance system software
- Experience working in environments where performance, scalability, and reliability are critical
- Familiarity with observability concepts: traces, logs, metrics, APM, and monitoring patterns
- Strong problem-solving and debugging skills with the ability to root-cause complex production issues
- Excellent communication skills, both written and verbal, with ability to collaborate in multicultural, remote-first teams
- Code quality mindset: you value simplicity, performance, maintainability, and thorough testing
- Experience with time-series data and understanding of metrics cardinality challenges
- Proficiency with SQL and experience working with relational or distributed databases
- Experience building cloud-native SaaS platforms with multi-tenant architecture
- Multi-cloud experience: working with AWS, GCP, Azure, or other cloud providers in a production setting
- Kubernetes proficiency: operating, monitoring, or developing for Kubernetes clusters
- Open-source observability tools: hands-on experience with Grafana, Alertmanager, Loki, Tempo, OpenTelemetry Collector, OTLP protocol
- OpenTelemetry expertise: experience with, or active contributions to OTel projects
- Time-series database experience: Prometheus TSDB, InfluxDB, Mimir, TimescaleDB, or SingleStore
- Experience with data pipeline technologies: Apache Kafka, Parquet, Arrow, or stream processing frameworks (Flink, etc.)
- Experience working with AI agents or LLM-powered applications, including agentic workflows for observability — enabling customers to query telemetry data in natural language
- Experience in a SaaS or cloud-native company delivering managed services to customers
Qualifications
Must Haves
- 2+ years of professional software development experience building distributed systems or backend services
- Strong proficiency in Go (Golang) — experience with Rust, Python, or C++ is also valuable
- Deep understanding of distributed systems concepts: scalability, consistency, high availability, concurrency, and failure modes
- Familiarity with distributed systems managed via Kubernetes
- Demonstrated ability to design and build reliable, high-performance system software
- Experience working in environments where performance, scalability, and reliability are critical
- Familiarity with observability concepts: traces, logs, metrics, APM, and monitoring patterns
- Strong problem-solving and debugging skills with the ability to root-cause complex production issues
- Excellent communication skills, both written and verbal, with ability to collaborate in multicultural, remote-first teams
- Code quality mindset: you value simplicity, performance, maintainability, and thorough testing
Nice to Haves
- Experience with time-series data and understanding of metrics cardinality challenges
- Proficiency with SQL and experience working with relational or distributed databases
- Experience building cloud-native SaaS platforms with multi-tenant architecture
- Multi-cloud experience: working with AWS, GCP, Azure, or other cloud providers in a production setting
- Kubernetes proficiency: operating, monitoring, or developing for Kubernetes clusters
- Open-source observability tools: hands-on experience with Grafana, Alertmanager, Loki, Tempo, OpenTelemetry Collector, OTLP protocol
- OpenTelemetry expertise: experience with, or active contributions to OTel projects
- Time-series database experience: Prometheus TSDB, InfluxDB, Mimir, TimescaleDB, or SingleStore
- Experience with data pipeline technologies: Apache Kafka, Parquet, Arrow, or stream processing frameworks (Flink, etc.)
- Experience working with AI agents or LLM-powered applications, including agentic workflows for observability — enabling customers to query telemetry data in natural language
- Experience in a SaaS or cloud-native company delivering managed services to customers
Benefits
- Certain roles are eligible for additional rewards, including merit increases and annual bonuses.
- Remote-first teams