Summary
Databento is a next-generation market data provider serving finance and fintech institutions with scalable, accessible financial data infrastructure. The Site Reliability Engineer will own platform uptime, performance, and observability while improving deployment, reliability, incident response, and operational practices across backend engineering.
Responsibilities
- Own uptime, SLAs, and SLOs across our API and platform services
- Set reliability and operational best practices for other developers without slowing them down
- Build and maintain observability across logging, metrics, and tracing
- Design and run high-availability deployment and containerization strategies
- Profile and optimize Python applications for throughput, latency, and cost
- Debug production issues down to the OS level using tools like strace, perf, eBPF, ss, and gdb
- Improve deployment and CI/CD workflows
- Join the on-call rotation, lead incident response, and run post-incident reviews
- Find what needs fixing on your own, then take projects from idea to completion
Skills
- Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup
- Hands-on experience with observability tooling for logging, metrics, and tracing (e.g. Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector)
- Experience with containerization and high availability deployment (e.g. Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s)
- Strong proficiency in Python, including application development and performance optimization
- Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb
- A track record of measurable impact in a recent role, such as improving performance by X%, speeding something up Nx, or saving $Y per year
- Experience with alerting and incident response best practices
- Familiarity with configuration management or infrastructure-as-code tools (Ansible, Terraform)
- HTTP benchmarking, load testing, and capacity planning
- Database schema design and query optimization skills
- Good communication skills and work ethic for a remote workplace
- An interest in financial data or algorithmic trading
Qualifications
Nice to Haves
- Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup
- Hands-on experience with observability tooling for logging, metrics, and tracing (e.g. Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector)
- Experience with containerization and high availability deployment (e.g. Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s)
- Strong proficiency in Python, including application development and performance optimization
- Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb
- A track record of measurable impact in a recent role, such as improving performance by X%, speeding something up Nx, or saving $Y per year
- Experience with alerting and incident response best practices
- Familiarity with configuration management or infrastructure-as-code tools (Ansible, Terraform)
- HTTP benchmarking, load testing, and capacity planning
- Database schema design and query optimization skills
- Good communication skills and work ethic for a remote workplace
- An interest in financial data or algorithmic trading
Benefits