Summary
TikTok is the leading destination for short-form mobile video, and they are seeking Software Engineering Interns to join the Model Infra team. The role focuses on optimizing the performance of recommendation systems, leveraging generative AI and large-scale model infrastructure.
Responsibilities
- Drive the optimization of training and inference pipelines to maximize hardware utilization (MFU/HFU) for models featuring hundreds of billions of dense parameters
- Architect specialized systems to support the integration of LLMs into the recommendation stack, focusing on memory-efficient attention mechanisms and advanced KV cache management for long-sequence user modeling
- Build and optimize high-concurrency engines for Petabyte-scale streaming training, handling continuous parameter updates and high-frequency data ingestion without compromising stability
- Work closely with researchers to design next-generation recommendation architectures optimized for modern GPU/NPU interconnects, ensuring high-bandwidth utilization across the cluster
- Innovate on how we store and synchronize massive model states across heterogeneous memory hierarchies (HBM, DDR, and NVMe)
Skills
- Currently pursuing an Undergraduate/Master in Software Development, Computer Science, Computer Engineering, or a related technical discipline
- Strong programming skills in C++ and Python
- Solid understanding of Computer Architecture and the GPU software stack (CUDA, Triton, or NCCL)
- Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and a desire to 'look under the hood' of model execution runtimes
- A strong interest in solving system-level bottlenecks in large-scale distributed environments
- Experience with Transformer-based architectures, 3D parallelism (TP/PP/DP)
- Deep understanding of the torch.compile stack, including TorchDynamo (graph acquisition) and TorchInductor (lowering)
- Hands-on experience writing high-performance kernels or optimizing collective communication (e.g., customizing NCCL/UCX)
- Familiarity with RDMA networking, high-performance storage, or specialized Parameter Server architectures
- Success in programming competitions (ACM-ICPC) or contributions to prominent open-source AI infrastructure or high-performance computing projects
Qualifications
Must Haves
- Currently pursuing an Undergraduate/Master in Software Development, Computer Science, Computer Engineering, or a related technical discipline
- Strong programming skills in C++ and Python
- Solid understanding of Computer Architecture and the GPU software stack (CUDA, Triton, or NCCL)
- Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and a desire to 'look under the hood' of model execution runtimes
- A strong interest in solving system-level bottlenecks in large-scale distributed environments
Nice to Haves
- Experience with Transformer-based architectures, 3D parallelism (TP/PP/DP)
- Deep understanding of the torch.compile stack, including TorchDynamo (graph acquisition) and TorchInductor (lowering)
- Hands-on experience writing high-performance kernels or optimizing collective communication (e.g., customizing NCCL/UCX)
- Familiarity with RDMA networking, high-performance storage, or specialized Parameter Server architectures
- Success in programming competitions (ACM-ICPC) or contributions to prominent open-source AI infrastructure or high-performance computing projects
Benefits
- Day one access to health insurance
- Life insurance
- Wellbeing benefits
- 10 paid holidays per year
- Paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year)
- Housing allowance