Databento logo
Databento
Posted 3 days agoVerified live 1d ago

Site Reliability Engineer

Brief overview

Remote
MastersOr in progress
3+ yrsMinimum
PythonSite Reliability EngineeringContainerizationHigh-Availability DeploymentLinux Debugging and ProfilingCI/CDIncident ResponsePerformance OptimizationTerraformAnsibleLoad Testing and Capacity Planning

About the company

Databento logo
Databentodatabento.com

Databento is a SaaS company that offers market data APIs for real-time and historical data.

Job description

Summary

Databento is a next-generation market data provider serving finance and fintech institutions with scalable, accessible financial data infrastructure. The Site Reliability Engineer will own platform uptime, performance, and observability while improving deployment, reliability, incident response, and operational practices across backend engineering.

Responsibilities

  • Own uptime, SLAs, and SLOs across our API and platform services
  • Set reliability and operational best practices for other developers without slowing them down
  • Build and maintain observability across logging, metrics, and tracing
  • Design and run high-availability deployment and containerization strategies
  • Profile and optimize Python applications for throughput, latency, and cost
  • Debug production issues down to the OS level using tools like strace, perf, eBPF, ss, and gdb
  • Improve deployment and CI/CD workflows
  • Join the on-call rotation, lead incident response, and run post-incident reviews
  • Find what needs fixing on your own, then take projects from idea to completion

Skills

  • Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup
  • Hands-on experience with observability tooling for logging, metrics, and tracing (e.g. Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector)
  • Experience with containerization and high availability deployment (e.g. Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s)
  • Strong proficiency in Python, including application development and performance optimization
  • Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb
  • A track record of measurable impact in a recent role, such as improving performance by X%, speeding something up Nx, or saving $Y per year
  • Experience with alerting and incident response best practices
  • Familiarity with configuration management or infrastructure-as-code tools (Ansible, Terraform)
  • HTTP benchmarking, load testing, and capacity planning
  • Database schema design and query optimization skills
  • Good communication skills and work ethic for a remote workplace
  • An interest in financial data or algorithmic trading

Qualifications

Nice to Haves

  • Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup
  • Hands-on experience with observability tooling for logging, metrics, and tracing (e.g. Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector)
  • Experience with containerization and high availability deployment (e.g. Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s)
  • Strong proficiency in Python, including application development and performance optimization
  • Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb
  • A track record of measurable impact in a recent role, such as improving performance by X%, speeding something up Nx, or saving $Y per year
  • Experience with alerting and incident response best practices
  • Familiarity with configuration management or infrastructure-as-code tools (Ansible, Terraform)
  • HTTP benchmarking, load testing, and capacity planning
  • Database schema design and query optimization skills
  • Good communication skills and work ethic for a remote workplace
  • An interest in financial data or algorithmic trading

Benefits

  • Remote workplace

More jobs like this