Summary
NVIDIA is a technology company building software and systems that power advanced generative AI and large language model workloads. The company is seeking a Software Engineer to bring up, benchmark, debug, analyze, and optimize distributed AI training and inference workloads across large-scale NVIDIA GPU platforms. The role also involves developing benchmarking tools, automation, debugging workflows, and resilience tooling for multi-GPU and multi-node environments.
Responsibilities
- Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads
- Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks
- Perform root-cause analysis of failures in large distributed environments
- Contribute to the resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster
- Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms
- Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams
- Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization
Skills
- Bachelor's or Master's in Computer Science or a related technical field (or equivalent experience)
- Experience developing software for AI, HPC, or systems-level applications
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution
- Background with debugging and scaling distributed systems
- Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware
- Experience operating workloads in scheduled, containerized cluster environments
- Excellent analytical, debugging, and communication skills, and a collaborative approach across teams
- Strong Python and C/C++ programming skills
- Hands-on experience with NCCL and CUDA-aware distributed execution
- Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging
- Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf
- Experience diagnosing performance jitter
- Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure
Qualifications
Must Haves
- Bachelor's or Master's in Computer Science or a related technical field (or equivalent experience)
- Experience developing software for AI, HPC, or systems-level applications
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution
- Background with debugging and scaling distributed systems
- Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware
- Experience operating workloads in scheduled, containerized cluster environments
- Excellent analytical, debugging, and communication skills, and a collaborative approach across teams
- Strong Python and C/C++ programming skills
Nice to Haves
- Hands-on experience with NCCL and CUDA-aware distributed execution
- Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging
- Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf
- Experience diagnosing performance jitter
- Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure
Benefits
- You will also be eligible for equity
- Benefits