TikTok USDS Joint Venture logo
TikTok USDS Joint Venture
Posted 84 days agoVerified live 2d ago

Site Reliability Engineer, Platform Responsibility - USDS

Brief overview

Seattle, WAIn-person
UndergradOr in progress
1+ yrsMinimum

Job description

Summary

TikTok USDS Joint Venture is a fast-growing company focused on building machine learning models to protect users from internet abuse and fraud. The Site Reliability Engineer will manage data service operations, design AI-powered automation, and improve system efficiency to enhance user experience.

Responsibilities

  • Manage day-to-day operations of data service, realtime/batch data pipelines, such as SLA/SLO/SLI management, system deployment, performance tuning and troubleshooting
  • Design and deploy AI Agents and LLM-powered automation to streamline incident response, root cause analysis, and proactive system monitoring
  • Create tools and automation to improve system administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality
  • Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement
  • Practice sustainable user support, incident response, and post mortem
  • This position is part of a team that provides 24/7 support and requires working scheduled shifts, which may include holidays

Skills

  • Bachelor or above degree in computer science or a related technical discipline
  • At least 1 year of industrial experience
  • Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
  • Demonstrated independent thinking capabilities and troubleshooting skills
  • Familiar with Unix/Linux system internals, networking, and distributed systems
  • Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches
  • In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
  • Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc

Qualifications

Must Haves

  • Bachelor or above degree in computer science or a related technical discipline
  • At least 1 year of industrial experience
  • Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
  • Demonstrated independent thinking capabilities and troubleshooting skills
  • Familiar with Unix/Linux system internals, networking, and distributed systems
  • Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches

Nice to Haves

  • In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
  • Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc

More jobs like this