NVIDIA logo
NVIDIA
Posted 50 days agoVerified live 1d ago

Senior System Software Engineer - GPU Performance

Brief overview

Remote
MastersOr in progress
$152k–$288k/yrStated range
3+ yrsMinimum
4,626 H-1B approvalsDept. of Labor
1,242 green cardsCertified filings
Performance EngineeringHPCParallel ProgrammingMPINCCLUCXNVSHMEMPerformance BenchmarkingComputer System ArchitectureC/C++PythonKubernetesSLURMAnsibleDockerCUDADeep Learning Frameworks

About the company

NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI.

Visa sponsorship history

4 years sponsoring, last filed FY2026H-1B dependent

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
4,626H-1B approved
99%approval rate
1,238new H-1B hires
1,242PERM certified
$190,000median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
2023997
20241,519
20251,767
2026343
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
2023203
2024276
2025466
2026436
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
2023362
2024279
2025573
202628
Top sponsored roles
Software EngineerHardware Engineer, ElectronicsEngineer Senior Systems SoftwareArchitectEngineer Senior ASIC
Sponsored employees from
IndiaChinaCanadaTaiwanSouth Korea

Job description

Summary

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The role involves conducting performance characterization and analysis on large multi-GPU and multi-node clusters, and collaborating with a dynamic team to enhance communication libraries for deep learning and HPC applications.

Responsibilities

  • Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters
  • Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack
  • Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available
  • Triage and root-cause performance issues reported by our customers
  • Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information
  • Collaborate with a very dynamic team across multiple time zones

Skills

  • M.S. (or equivalent experience) or PhD in Computer Science, or related field with relevant performance engineering and HPC experience
  • 3+ yrs of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
  • Experience conducting performance benchmarking and triage on large scale HPC clusters
  • Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
  • Implement micro-benchmarks in C/C++, read and modify the code base when required
  • Ability to debug performance issues across the entire HW/SW stack. Proficient in a scripting language, preferably Python
  • Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible, Docker)
  • Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively across different teams and timezones
  • Practical experience with Infiniband/Ethernet networks in areas like RDMA, topologies, congestion control
  • Experience debugging network issues in large scale deployments
  • Familiarity with CUDA programming and/or GPUs
  • Experience with Deep Learning Frameworks such PyTorch, TensorFlow

Qualifications

Must Haves

  • M.S. (or equivalent experience) or PhD in Computer Science, or related field with relevant performance engineering and HPC experience
  • 3+ yrs of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
  • Experience conducting performance benchmarking and triage on large scale HPC clusters
  • Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
  • Implement micro-benchmarks in C/C++, read and modify the code base when required
  • Ability to debug performance issues across the entire HW/SW stack. Proficient in a scripting language, preferably Python
  • Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible, Docker)
  • Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively across different teams and timezones

Nice to Haves

  • Practical experience with Infiniband/Ethernet networks in areas like RDMA, topologies, congestion control
  • Experience debugging network issues in large scale deployments
  • Familiarity with CUDA programming and/or GPUs
  • Experience with Deep Learning Frameworks such PyTorch, TensorFlow

Benefits

  • Equity
  • Benefits

More jobs like this