Summary
TikTok USDS Joint Venture is a fast-growing company focused on building machine learning models to protect users from internet abuse and fraud. The Site Reliability Engineer will manage data service operations, design AI-powered automation, and improve system efficiency to enhance user experience.
Responsibilities
- Manage day-to-day operations of data service, realtime/batch data pipelines, such as SLA/SLO/SLI management, system deployment, performance tuning and troubleshooting
- Design and deploy AI Agents and LLM-powered automation to streamline incident response, root cause analysis, and proactive system monitoring
- Create tools and automation to improve system administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality
- Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement
- Practice sustainable user support, incident response, and post mortem
- This position is part of a team that provides 24/7 support and requires working scheduled shifts, which may include holidays
Skills
- Bachelor or above degree in computer science or a related technical discipline
- At least 1 year of industrial experience
- Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
- Demonstrated independent thinking capabilities and troubleshooting skills
- Familiar with Unix/Linux system internals, networking, and distributed systems
- Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches
- In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
- Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc
Qualifications
Must Haves
- Bachelor or above degree in computer science or a related technical discipline
- At least 1 year of industrial experience
- Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
- Demonstrated independent thinking capabilities and troubleshooting skills
- Familiar with Unix/Linux system internals, networking, and distributed systems
- Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches
Nice to Haves
- In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
- Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc