Summary
Tyk provides an API Management platform that helps organizations connect systems and services. The Site Reliability Engineer will manage, maintain, improve, and support the Tyk Cloud platform, while identifying reliability improvements, automating operations, managing incidents, and contributing to multi-region and multi-cloud infrastructure.
Responsibilities
- Maintaining global Tyk Cloud within SL(A/I/O)s you will help to define
- Identifying reliability issues and working together with your squad to solve them
- Identifying and introducing new metrics and building relevant dashboards
- Participating in the on-call rotation
- Working with your squad to expand multi-region and multi-cloud reach of the platform
- Documenting operational knowledge
- Conducting post-incident analysis
- Automating common tasks
- Be a key shaper and contributor to our continuous improvement agenda – be it the clarity of our user stories, how we estimate, communicate with other teams or customers – we expect this role to be advocate of continuous improvement
- Reliability of our new global Tyk Cloud platform
- Automation of operations and support
- Writing and maintaining documentation on SRE processes and policies
- Recommending and implementing ways of driving operational efficiency and driving down our cost to run, without impacting service
- Assisting in penetration testing for Cloud through liaising with our provider, providing technical details, and environment setup
- Incident management
Skills
- * Strong collaboration skills
- * Launching and operating production scale kubernetes clusters
- * Designing and operating infrastructure on AWS and other providers
- * Operating MongoDB (or other document database) clusters
- * Operating Redis (or other key-value storage) clusters
- * Administering Linux servers
- * Maintaining distributed software
- * Operating Prometheus and Grafana
- * Operating logging collection and analysis systems
- * Participating in the on-call rotation(16:00pm – 4:00am UTC)
- * Kubernetes & containers (advanced)
- * AWS / EKS (advanced)
- * Linux (advanced)
- * Terraform and IaC in general (proficient)
- * Helm (proficient)
- * Go (familiar)
- * MongoDB (or similar)
- * Redis (or similar)
- * Monitoring – prometheus, grafana, thanos (familiar)
- * Grasp of networking concepts (subnets, routing, peering, load balancing, NAT, etc.)
- * Common networking protocols (DNS, TCP/IP, HTTP, TLS, UDP)
- * Proactive, energetic, innovative and change oriented
- * GCP or Azure
- * Bare metal infrastructure engineering
- * API management experience
- * Large scale distributed storage management
- * Familiarity with Rancher
- * CKA/CKAD/CKS
- * Creating and delivering production software in Go language
Qualifications
Must Haves
- * Strong collaboration skills
- * Launching and operating production scale kubernetes clusters
- * Designing and operating infrastructure on AWS and other providers
- * Operating MongoDB (or other document database) clusters
- * Operating Redis (or other key-value storage) clusters
- * Administering Linux servers
- * Maintaining distributed software
- * Operating Prometheus and Grafana
- * Operating logging collection and analysis systems
- * Participating in the on-call rotation(16:00pm – 4:00am UTC)
- * Kubernetes & containers (advanced)
- * AWS / EKS (advanced)
- * Linux (advanced)
- * Terraform and IaC in general (proficient)
- * Helm (proficient)
- * Go (familiar)
- * MongoDB (or similar)
- * Redis (or similar)
- * Monitoring – prometheus, grafana, thanos (familiar)
- * Grasp of networking concepts (subnets, routing, peering, load balancing, NAT, etc.)
- * Common networking protocols (DNS, TCP/IP, HTTP, TLS, UDP)
- * Proactive, energetic, innovative and change oriented
Nice to Haves
- * GCP or Azure
- * Bare metal infrastructure engineering
- * API management experience
- * Large scale distributed storage management
- * Familiarity with Rancher
- * CKA/CKAD/CKS
- * Creating and delivering production software in Go language
Benefits
- Unlimited paid holidays
- Remote working from anywhere in the world
- Total flexibility in hours
- Employee share scheme
- Generous maternity and paternity leave
- Company retreats