TikTok USDS Joint Venture logo
TikTok USDS Joint Venture
Posted 84 days agoVerified live 22h ago

Site Reliability Engineer, Platform Responsibility - USDS

Brief overview

San Jose, CAIn-person
UndergradOr in progress
1+ yrsMinimum

Job description

Summary

TikTok USDS Joint Venture is a fast-growing company responsible for building machine learning models and systems to identify and defend against internet abuse and fraud. They are seeking a Site Reliability Engineer to manage data services, design automation for incident response, and improve operational efficiency.

Responsibilities

  • Manage day-to-day operations of data service, realtime/batch data pipelines, such as SLA/SLO/SLI management, system deployment, performance tuning and troubleshooting
  • Design and deploy AI Agents and LLM-powered automation to streamline incident response, root cause analysis, and proactive system monitoring
  • Create tools and automation to improve system administration and operational efficiency, leveraging AI-assisted development tools to accelerate delivery and code quality
  • Participate in regular on-call rotations as part of a team that provides 24 hour coverage across multiple shifts
  • Engage in and improve the whole lifecycle of services from inception and design, development, capacity planning, and launch reviews, to deployment, operation, and refinement
  • Practice sustainable user support, incident response, and post mortem

Skills

  • Bachelor or above degree in computer science or a related technical discipline
  • At least 1 year of industrial experience
  • Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
  • Demonstrated independent thinking capabilities and troubleshooting skills
  • Familiar with Unix/Linux system internals, networking, and distributed systems
  • Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches
  • In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
  • Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc

Qualifications

Must Haves

  • Bachelor or above degree in computer science or a related technical discipline
  • At least 1 year of industrial experience
  • Experience integrating AI/LLM APIs into internal workflows or infrastructure tooling
  • Demonstrated independent thinking capabilities and troubleshooting skills
  • Familiar with Unix/Linux system internals, networking, and distributed systems
  • Expertise in monitoring tools (e.g., Prometheus, Grafana, DataDog) and fundamental observability approaches

Nice to Haves

  • In-depth knowledge of Unix/Linux systems, networking fundamentals and system performance tuning
  • Familiar with backend systems such as MySQL/Redis/Nginx/Kafka/Kubernetes/Docker and big data technologies such as Hadoop/Spark/Flink/Hive/OLAP/ClickHouse, etc

More jobs like this