Summary
ByteDance is a rapidly growing company focused on inspiring creativity and enriching life. They are seeking a Software Engineer to join their Inference Infrastructure team, responsible for designing and operating platforms for AI workloads and large-scale LLM inference.
Responsibilities
- Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience
- Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms
- Collaborate across teams to deliver world-class inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines
- Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems
- Write high-quality, production-ready code that is maintainable, testable, and scalable
Skills
- B.S./M.S. in Computer Science, Computer Engineering, or related fields with 2+ years of relevant experience (Ph.D. with strong systems/ML publications also considered)
- Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems
- Hands-on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration
- Solid knowledge of container and orchestration technologies (Docker, Kubernetes)
- Proficiency in at least one major programming language (Go, Rust, Python, or C++)
- Experience contributing to or operating large-scale cluster management systems (e.g., Kubernetes, Ray)
- Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments
- Hands-on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT-LLM)
- Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI)
- Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms
- Excellent communication skills and ability to collaborate across global, cross-functional teams
- Passion for system efficiency, performance optimization, and open-source innovation
Qualifications
Must Haves
- B.S./M.S. in Computer Science, Computer Engineering, or related fields with 2+ years of relevant experience (Ph.D. with strong systems/ML publications also considered)
- Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems
- Hands-on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration
- Solid knowledge of container and orchestration technologies (Docker, Kubernetes)
- Proficiency in at least one major programming language (Go, Rust, Python, or C++)
Nice to Haves
- Experience contributing to or operating large-scale cluster management systems (e.g., Kubernetes, Ray)
- Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments
- Hands-on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT-LLM)
- Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI)
- Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms
- Excellent communication skills and ability to collaborate across global, cross-functional teams
- Passion for system efficiency, performance optimization, and open-source innovation
Benefits
- Employees have day one access to medical, dental, and vision insurance
- A 401(k) savings plan with company match
- Paid parental leave
- Short-term and long-term disability coverage
- Life insurance
- Wellbeing benefits
- Employees also receive 10 paid holidays per year
- 10 paid sick days per year
- 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure)