Modular, a Qualcomm company logo
Modular, a Qualcomm company
Posted 18 days agoVerified live 22h ago

Driver Engineer

Brief overview

Remote
MastersOr in progress
$148k–$270k/yrStated range
3+ yrsMinimum
C++17/20CUDA Driver APIHIPMetalVulkan ComputeAccelerator Execution and Memory ModelsConcurrency and ParallelismLibrary and API DesignGDB/LLDBSanitizers

Job description

Summary

Modular, a Qualcomm company, develops compiler and runtime technologies that enable software to run across diverse accelerators. The Driver Engineer will build and extend core driver abstractions, communication primitives, diagnostics, and hardware backend integrations supporting MAX and Mojo across multi-accelerator systems.

Responsibilities

  • Implement and extend core driver abstractions — Device, Context, Queue, Memory, Function (kernel representation and launch) — across diverse hardware backends, including bindings that expose vendor-specific functionality where it matters
  • Contribute to multi-accelerator and multi-node communication and collectives primitives that underpin large-model inference, using technologies like NVLink, RDMA (Infiniband, RoCE, EFA), and sockets — via a combination of direct programming and broker libraries like UCX
  • Improve diagnostics and error reporting across our asynchronous execution stack — turning low-level driver failures into actionable messages for kernel and graph authors in production environments
  • Partner with peer teams — Kernels, Graph Compiler/Runtime, and Serving — to build and refine the surfaces they consume in the Mojo standard library and MAX framework
  • Participate in design discussions, code reviews, and collaborative software development to uphold high engineering standards
  • Contribute to a fully open source project — everything you build will be public and part of our GitHub repo

Skills

  • * 3+ years of experience writing high-performance, low-latency production systems C++ (modern C++17/20), with strong instincts for ownership, lifetime, ABI stability, concurrency, and parallelism
  • * Hands-on experience with at least one accelerator driver-level API — CUDA Driver API, HIP, Metal, or Vulkan compute — including streams, events, contexts, and module loading. Familiarity with async allocators and IPC is a plus
  • * Working understanding of accelerator execution and memory models — stream ordering, host↔device transfers, and pinned memory. Exposure to peer access and multi-GPU topologies (NVLink/xGMI, NUMA) is a plus
  • * Solid instincts for library and API design — naming, layering, ergonomics
  • * Strong debugging skills in asynchronous, multi-device systems — GDB/LLDB, sanitizers, CUDA/HIP error modes, race conditions, leaked resources
  • * A proactive, collaborative mindset, intellectual curiosity, and drive to identify and solve complex technical challenges as part of a high-performing team
  • Candidates based in the US and Canada are welcome to apply
  • Those in earlier career stages work in a hybrid capacity at our Los Altos, CA office (minimum 3 days per week on-site) with relocation assistance provided for out-of-state candidates based in the US
  • Senior members have both in office or remote flexibility
  • * Experience with asynchronous runtimes, custom memory allocators, or RDMA-based networking (Infiniband, RoCE, EFA, UCX, NIXL, MPI, NVSHMEM/ROCSHMEM/OpenSHMEM)
  • * Exposure to non-NVIDIA accelerators in production — AMD ROCm, Apple Metal, etc
  • * Experience with zero-copy tensor interoperability — DLPack, the CUDA Array Interface, or similar
  • * Familiarity with Mojo, or recent open-source contributions to systems projects (LLVM, PyTorch, JAX/XLA, TVM, IREE, vLLM, TRT-LLM)
  • * An advanced degree in Computer Science or a related field

Qualifications

Must Haves

  • * 3+ years of experience writing high-performance, low-latency production systems C++ (modern C++17/20), with strong instincts for ownership, lifetime, ABI stability, concurrency, and parallelism
  • * Hands-on experience with at least one accelerator driver-level API — CUDA Driver API, HIP, Metal, or Vulkan compute — including streams, events, contexts, and module loading. Familiarity with async allocators and IPC is a plus
  • * Working understanding of accelerator execution and memory models — stream ordering, host↔device transfers, and pinned memory. Exposure to peer access and multi-GPU topologies (NVLink/xGMI, NUMA) is a plus
  • * Solid instincts for library and API design — naming, layering, ergonomics
  • * Strong debugging skills in asynchronous, multi-device systems — GDB/LLDB, sanitizers, CUDA/HIP error modes, race conditions, leaked resources
  • * A proactive, collaborative mindset, intellectual curiosity, and drive to identify and solve complex technical challenges as part of a high-performing team
  • Candidates based in the US and Canada are welcome to apply
  • those in earlier career stages work in a hybrid capacity at our Los Altos, CA office (minimum 3 days per week on-site) with relocation assistance provided for out-of-state candidates based in the US
  • Senior members have both in office or remote flexibility

Nice to Haves

  • * Experience with asynchronous runtimes, custom memory allocators, or RDMA-based networking (Infiniband, RoCE, EFA, UCX, NIXL, MPI, NVSHMEM/ROCSHMEM/OpenSHMEM)
  • * Exposure to non-NVIDIA accelerators in production — AMD ROCm, Apple Metal, etc
  • * Experience with zero-copy tensor interoperability — DLPack, the CUDA Array Interface, or similar
  • * Familiarity with Mojo, or recent open-source contributions to systems projects (LLVM, PyTorch, JAX/XLA, TVM, IREE, vLLM, TRT-LLM)
  • * An advanced degree in Computer Science or a related field

Benefits

  • Comprehensive healthcare coverage
  • Retirement and savings programs
  • Employee stock purchase opportunities
  • Paid time off
  • Wellbeing resources
  • Family support programs
  • Learning and development opportunities
  • RSU grants
  • Annual target bonus
  • Equity
  • Relocation assistance for out-of-state candidates based in the US
  • Senior members have both in office or remote flexibility
  • Regular team onsites and local meetups in Los Altos, CA as well as different cities

More jobs like this