Summary
Ookla, an Accenture company, provides connectivity intelligence and network performance insights through services such as Speedtest, Downdetector, Ekahau, and RootMetrics. The Site Reliability Engineer will build, maintain, and operate large-scale infrastructure supporting Ookla services, with responsibilities spanning reliability, scalability, security, observability, databases, cloud platforms, deployments, and operational support.
Responsibilities
- Maintaining a distributed, global ecosystem of thousands of cloud instances, containerized workflows, serverless applications, Linux servers, and associated infrastructure supporting billions of requests daily
- Maintaining transactional database infrastructure using MySQL, PostgreSQL, and managed services such as RDS/Aurora
- Supporting the use of NoSQL data storage engines such as DynamoDb and MongoDB
- Building and supporting data stream processing with Kinesis or Kafka
- Supporting data engineering and big data toolchains such as Spark
- Supporting production systems in a 24x7x365 environment, including on-call responsibilities
- Providing architectural and operational support to software engineers in a wide variety of focus areas
- Support software and data engineering teams and guiding operational best practices
- Implementation and oversight of security programs including vulnerability remediation, patch management, IDS/IPS, penetration testing, and interfacing with our corporate InfoSec team
- Supporting the development to production code deploy pipeline for a range of production applications
- Providing the tooling and guidance for the software and data engineering team to implement our monitoring and observability best practices
- Assisting development teams with troubleshooting
Skills
- 4+ Years Systems/Platform engineering experience
- Experience building globally-distributed systems
- Strong understanding of security best practices
- Infrastructure as Code: Terraform, Cloudformation
- Branching and Merge based Source Code Configuration Management: Git, Github
- Configuration management systems such as Chef or Ansible
- Container-based architectures including Docker, Kubernetes
- Proficiency in one or more high level programming languages such as Typescript, Go, Python, PHP, Ruby, Java, etc
- Experience with AWS and other Cloud infrastructure platforms
- Comfort writing SQL queries and analyzing query performance
- Comfortable learning and working with new technologies in an ever-changing environment
- Strong verbal and written communication skills
- Strong time management skills and a self-driven work ethic
Qualifications
Must Haves
- 4+ Years Systems/Platform engineering experience
- Experience building globally-distributed systems
- Strong understanding of security best practices
- Infrastructure as Code: Terraform, Cloudformation
- Branching and Merge based Source Code Configuration Management: Git, Github
- Configuration management systems such as Chef or Ansible
- Container-based architectures including Docker, Kubernetes
- Proficiency in one or more high level programming languages such as Typescript, Go, Python, PHP, Ruby, Java, etc
- Experience with AWS and other Cloud infrastructure platforms
- Comfort writing SQL queries and analyzing query performance
- Comfortable learning and working with new technologies in an ever-changing environment
- Strong verbal and written communication skills
- Strong time management skills and a self-driven work ethic
Benefits
- A flexible work environment where individuality, fun, and talent are all valued equally.