Summary
TikTok is a company dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love. They are seeking an experienced Site Reliability Engineer to improve the Lark system, ensuring software reliability and managing production systems.
Responsibilities
- Responsible for overall reliability of Lark product
- Perform lifecycle management of production systems including change management, service deployment, operations and emergency response
- Monitor the system and respond to incidents to maintain system service level agreement (SLA), review and follow up all production incidents
- Perform capacity management of compute, storage and network bandwidth resources to ensure system stability and save infrastructure costs
- Provide strong support during big events to ensure the system is capable of consuming a large volume of Internet traffic
- Build tools, automations, visualizations and monitors to facilitate the operation and optimization of the global infrastructure
Skills
- Bachelor's degree in Computer Science or a related technical background involving software/system engineering, or equivalent working experience
- 1+ years of SRE or DevOps experience in large scale online services
- Programming experience with at least one of the following languages: C, C++, Java, Python, C# or Go
- Extensive knowledge of networking, operation systems, database systems and container technology
- Good understanding of every aspect of microservice architecture, and hands on experience in troubleshooting in large scale distributed systems
- Hands on experience in common opensource systems such as Linux, MySQL, MongoDB, Redis and ELK
- Experience in building solutions with AWS, Google, Azures and other cloud services is a plus
- Passionate, self-motivated and good teamwork skills
Qualifications
Must Haves
- Bachelor's degree in Computer Science or a related technical background involving software/system engineering, or equivalent working experience
- 1+ years of SRE or DevOps experience in large scale online services
- Programming experience with at least one of the following languages: C, C++, Java, Python, C# or Go
Nice to Haves
- Extensive knowledge of networking, operation systems, database systems and container technology
- Good understanding of every aspect of microservice architecture, and hands on experience in troubleshooting in large scale distributed systems
- Hands on experience in common opensource systems such as Linux, MySQL, MongoDB, Redis and ELK
- Experience in building solutions with AWS, Google, Azures and other cloud services is a plus
- Passionate, self-motivated and good teamwork skills
Benefits
- Employees have day one access to medical, dental, and vision insurance
- A 401(k) savings plan with company match
- Paid parental leave
- Short-term and long-term disability coverage
- Life insurance
- Wellbeing benefits
- 10 paid holidays per year
- 10 paid sick days per year
- 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure)