Arango logo
Arango
Posted 17 days agoVerified live 4d ago

Site Reliability Engineer (Europe)

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
AWSGoogle Cloud Platform (GCP)Linux InternalsDockerKubernetesCI/CD PipelinesJenkinsCircleCIPrometheusGrafanaELK StackGitGolangPythonCore NetworkingDistributed DatabasesAutonomous teamwork

About the company

Arango logo
Arangoarango.ai

Arango provides the trusted data foundation for enterprise AI through its Contextual Data Platform, transforming fragmented enterprise data into a contextual data layer that enables AI systems to operate with business context at scale.

Job description

Summary

Arango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with trusted business context. The Site Reliability Engineer will maintain and improve the reliability, scalability, and performance of distributed database systems running on Kubernetes and cloud environments, while automating infrastructure, enhancing observability, optimizing CI/CD pipelines, and troubleshooting complex system issues.

Responsibilities

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms
  • Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems
  • Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment
  • Develop strategies for disaster recovery, high availability, and fault tolerance
  • Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure)
  • Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance
  • Participate in on-call rotations to support critical production systems and respond to incidents
  • Collaborate with cross-functional teams to improve overall system reliability and scalability
  • Collaborate with the Customer Success team to resolve customer issues

Skills

  • • Experience: SRE or DevOps Engineer background in cloud-native environments. Self-organized, autonomous remote team player with strong communication skills
  • • Cloud & Infrastructure: AWS and GCP; advanced Linux internals (processes, environment variables); containerization and orchestration (Docker, Kubernetes at scale)
  • • CI/CD & Observability: CI/CD pipelines (Jenkins, CircleCI); monitoring, alerting, and logging (Prometheus, Grafana, ELK stack); Git version control
  • • Networking & Security: Core networking, security best practices, and systematic troubleshooting of complex infrastructure issues
  • • Development: Programming proficiency in Golang or Python
  • • Experience managing distributed databases or large-scale data storage systems
  • • Knowledge of security best practices in cloud environments
  • • Experience with scripting languages like Python or Bash
  • • Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus
  • • Experience working with GitOps
  • • Strong programming skills in Golang, with experience in developing automation tools, scripts, or services

Qualifications

Must Haves

  • • Experience: SRE or DevOps Engineer background in cloud-native environments. Self-organized, autonomous remote team player with strong communication skills
  • • Cloud & Infrastructure: AWS and GCP; advanced Linux internals (processes, environment variables); containerization and orchestration (Docker, Kubernetes at scale)
  • • CI/CD & Observability: CI/CD pipelines (Jenkins, CircleCI); monitoring, alerting, and logging (Prometheus, Grafana, ELK stack); Git version control
  • • Networking & Security: Core networking, security best practices, and systematic troubleshooting of complex infrastructure issues
  • • Development: Programming proficiency in Golang or Python

Nice to Haves

  • • Experience managing distributed databases or large-scale data storage systems
  • • Knowledge of security best practices in cloud environments
  • • Experience with scripting languages like Python or Bash
  • • Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus
  • • Experience working with GitOps
  • • Strong programming skills in Golang, with experience in developing automation tools, scripts, or services

Benefits

  • Remote work in India

More jobs like this