Summary
OVA.Work is seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications. The ideal candidate will combine software engineering and operations expertise to improve system reliability, automate operational processes, optimize performance, and ensure service availability.
Responsibilities
- Design, implement, and maintain highly available and scalable infrastructure for production environments
- Automate infrastructure provisioning, deployment, monitoring, and operational workflows
- Improve system reliability, availability, scalability, and performance through engineering best practices
- Develop and maintain CI/CD pipelines to support automated application deployments
- Monitor production systems, troubleshoot incidents, perform root cause analysis (RCA), and implement preventive measures
- Configure and manage observability solutions including logging, monitoring, tracing, and alerting
- Collaborate with development teams to improve application reliability and operational readiness
- Implement disaster recovery, backup, and business continuity strategies
- Optimize cloud infrastructure utilization and operational costs
- Develop automation scripts and tools to eliminate repetitive operational tasks
- Participate in incident response, on-call rotations, and post-incident reviews
- Document infrastructure architecture, operational procedures, and system configurations
Skills
- Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related field
- 3–6 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering
- Strong proficiency in Linux system administration
- Experience with Python, Go, Bash, or Shell scripting
- Hands-on experience with Docker and Kubernetes
- Strong understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, and load balancing
- Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps
- Familiarity with cloud platforms such as AWS, Microsoft Azure, or Google Cloud
- Strong understanding of infrastructure automation and configuration management
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, Pulumi, or CloudFormation
- Hands-on experience with monitoring and observability tools including Prometheus, Grafana, ELK Stack, OpenTelemetry, Datadog, or Splunk
- Familiarity with service mesh technologies such as Istio or Linkerd
- Experience managing distributed systems and microservices architectures
- Knowledge of security best practices, identity management, and compliance requirements
- Experience supporting high-availability, mission-critical production systems
- Experience supporting AI/ML infrastructure or MLOps platforms
- Familiarity with container security and Kubernetes security best practices
- Knowledge of FinOps and cloud cost optimization
- Experience with chaos engineering and resilience testing
- Relevant certifications such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), Google Professional Cloud DevOps Engineer, or Microsoft Azure DevOps Engineer Expert
Qualifications
Must Haves
- Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related field
- 3–6 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering
- Strong proficiency in Linux system administration
- Experience with Python, Go, Bash, or Shell scripting
- Hands-on experience with Docker and Kubernetes
- Strong understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, and load balancing
- Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps
- Familiarity with cloud platforms such as AWS, Microsoft Azure, or Google Cloud
- Strong understanding of infrastructure automation and configuration management
Nice to Haves
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, Pulumi, or CloudFormation
- Hands-on experience with monitoring and observability tools including Prometheus, Grafana, ELK Stack, OpenTelemetry, Datadog, or Splunk
- Familiarity with service mesh technologies such as Istio or Linkerd
- Experience managing distributed systems and microservices architectures
- Knowledge of security best practices, identity management, and compliance requirements
- Experience supporting high-availability, mission-critical production systems
- Experience supporting AI/ML infrastructure or MLOps platforms
- Familiarity with container security and Kubernetes security best practices
- Knowledge of FinOps and cloud cost optimization
- Experience with chaos engineering and resilience testing
- Relevant certifications such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), Google Professional Cloud DevOps Engineer, or Microsoft Azure DevOps Engineer Expert
Benefits
- Competitive salary and performance-based incentives.
- Comprehensive health and wellness benefits.
- Flexible or hybrid work arrangements.
- Learning, certification, and conference sponsorship opportunities.
- Access to modern cloud infrastructure and engineering tools.
- Opportunity to work on highly scalable, mission-critical systems in a collaborative and innovative environment.