ByteDance logo
ByteDance
Posted 142 days agoVerified live 2d ago

Software Engineer - AI Compute Infrastructure

Brief overview

Seattle, Washington, United States of AmericaIn-person
UndergradOr in progress
$148k–$301k/yrStated range
2+ yrsMinimum
2,631 H-1B approvalsDept. of Labor
463 green cardsCertified filings
Large model inferenceDistributed systemsParallel systemsHigh-performance networking systemsResource managementSchedulingRequest routingMonitoringOrchestrationDockerKubernetesGoRustPythonC++GPU programmingCUDA

About the company

ByteDance logo
ByteDancebytedance.com

ByteDance is a technology company that develops content creation platforms and services.

Visa sponsorship history

4 years sponsoring, last filed FY2026H-1B dependent

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
2,631H-1B approved
99%approval rate
1,036new H-1B hires
463PERM certified
$204,340median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
2023605
2024997
2025933
202696
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
2023234
2024196
2025193
202680
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
2023113
2024140
2025157
202653
Top sponsored roles
Software EngineerProduct ManagerData ScientistBackend Software EngineerResearch Scientist
Sponsored employees from
ChinaIndiaCanadaTaiwanHong Kong

Job description

Summary

ByteDance is a rapidly growing company focused on inspiring creativity and enriching life. They are seeking a Software Engineer to join their Inference Infrastructure team, responsible for designing and operating platforms for AI workloads and large-scale LLM inference.

Responsibilities

  • Design and build large-scale, container-based cluster management and orchestration systems with extreme performance, scalability, and resilience
  • Architect next-generation cloud-native GPU and AI accelerator infrastructure to deliver cost-efficient and secure ML platforms
  • Collaborate across teams to deliver world-class inference solutions using vLLM, SGLang, TensorRT-LLM, and other LLM engines
  • Stay current with the latest advances in open source (Kubernetes, Ray, etc.), AI/ML and LLM infrastructure, and systems research; integrate best practices into production systems
  • Write high-quality, production-ready code that is maintainable, testable, and scalable

Skills

  • B.S./M.S. in Computer Science, Computer Engineering, or related fields with 2+ years of relevant experience (Ph.D. with strong systems/ML publications also considered)
  • Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems
  • Hands-on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes)
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++)
  • Experience contributing to or operating large-scale cluster management systems (e.g., Kubernetes, Ray)
  • Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments
  • Hands-on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT-LLM)
  • Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI)
  • Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms
  • Excellent communication skills and ability to collaborate across global, cross-functional teams
  • Passion for system efficiency, performance optimization, and open-source innovation

Qualifications

Must Haves

  • B.S./M.S. in Computer Science, Computer Engineering, or related fields with 2+ years of relevant experience (Ph.D. with strong systems/ML publications also considered)
  • Strong understanding of large model inference, distributed and parallel systems, and/or high-performance networking systems
  • Hands-on experience building cloud or ML infrastructure in areas such as resource management, scheduling, request routing, monitoring, or orchestration
  • Solid knowledge of container and orchestration technologies (Docker, Kubernetes)
  • Proficiency in at least one major programming language (Go, Rust, Python, or C++)

Nice to Haves

  • Experience contributing to or operating large-scale cluster management systems (e.g., Kubernetes, Ray)
  • Experience with workload scheduling, GPU orchestration, scaling, and isolation in production environments
  • Hands-on experience with GPU programming (CUDA) or inference engines (vLLM, SGLang, TensorRT-LLM)
  • Familiarity with public cloud providers (AWS, Azure, GCP) and their ML platforms (SageMaker, Azure ML, Vertex AI)
  • Strong knowledge of ML systems (Ray, DeepSpeed, PyTorch) and distributed training/inference platforms
  • Excellent communication skills and ability to collaborate across global, cross-functional teams
  • Passion for system efficiency, performance optimization, and open-source innovation

Benefits

  • Employees have day one access to medical, dental, and vision insurance
  • A 401(k) savings plan with company match
  • Paid parental leave
  • Short-term and long-term disability coverage
  • Life insurance
  • Wellbeing benefits
  • Employees also receive 10 paid holidays per year
  • 10 paid sick days per year
  • 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure)

More jobs like this