Modular, a Qualcomm company logo
Modular, a Qualcomm company
Posted 18 days agoVerified live 1d ago

Mojo Libraries Engineer

Brief overview

Remote
UndergradOr in progress
$148k–$270k/yrStated range
2+ yrsMinimum
C++High-Performance ComputingCompiler EngineeringHeterogeneous Programming ModelsCUDASYCLOpenCLAccelerator ArchitecturesGPU Kernel DevelopmentPyTorch C++ APITritonCUTLASSCuTeMLIRLLVMHardware Platform Bring-UpModel Serving

Job description

Summary

Modular, a Qualcomm company, is building an AI platform and unified compute layer for developing and deploying AI models across diverse hardware. The Mojo Libraries Engineer will join the Hardware Enablement team to implement and optimize support for new accelerator architectures across the Modular software stack, collaborating with internal teams and external hardware partners.

Responsibilities

  • Implement and validate support for new hardware architectures across the Modular stack, working under the guidance of senior engineers on the team
  • Write and optimize Mojo kernels targeting novel accelerator architectures, with a focus on correctness first and performance iteration
  • Contribute to cross-team efforts improving portability infrastructure, tooling, and debugging workflows for new target hardware
  • Collaborate with hardware vendor engineers to understand target platforms, build integration tests, and triage platform-specific issues
  • Develop working knowledge of new hardware platforms — including ISA documentation, memory hierarchies, and vendor toolchains — and share findings with the team through demos and write-ups
  • Participate in company events such as on-sites and hackathons, contributing to a collaborative and open engineering culture

Skills

  • Candidates based in the US, Canada are welcome to apply
  • 2+ years of experience in high-performance computing, compiler engineering, or related domains in industry or research
  • Proficiency in C++ and experience working in complex, multi-component software systems
  • Hands-on experience with at least one heterogeneous programming model (CUDA, SYCL, OpenCL, or similar), either as a user or contributor
  • Some exposure to non-GPU accelerator architectures (DSPs, NPUs, or other hardware accelerators) is a strong plus
  • Curiosity and willingness to learn new hardware platforms quickly, comfortable reading architecture manuals and vendor documentation
  • A collaborative, team-oriented attitude and alignment with our culture
  • Familiarity with how AI operators are implemented at a low level (e.g., experience writing or modifying GPU kernels, custom operators, or working with frameworks like PyTorch at the C++ layer)
  • Experience with GPU DSLs/DSELs such as Triton, CUTLASS, or CuTe
  • Familiarity with MLIR or LLVM compiler infrastructure
  • Experience working directly with hardware vendor teams or on platform bring-up efforts
  • Exposure to model serving or inference optimization workflows

Qualifications

Must Haves

  • Candidates based in the US, Canada are welcome to apply
  • 2+ years of experience in high-performance computing, compiler engineering, or related domains in industry or research
  • Proficiency in C++ and experience working in complex, multi-component software systems
  • Hands-on experience with at least one heterogeneous programming model (CUDA, SYCL, OpenCL, or similar), either as a user or contributor
  • Some exposure to non-GPU accelerator architectures (DSPs, NPUs, or other hardware accelerators) is a strong plus
  • Curiosity and willingness to learn new hardware platforms quickly, comfortable reading architecture manuals and vendor documentation
  • A collaborative, team-oriented attitude and alignment with our culture

Nice to Haves

  • Familiarity with how AI operators are implemented at a low level (e.g., experience writing or modifying GPU kernels, custom operators, or working with frameworks like PyTorch at the C++ layer)
  • Experience with GPU DSLs/DSELs such as Triton, CUTLASS, or CuTe
  • Familiarity with MLIR or LLVM compiler infrastructure
  • Experience working directly with hardware vendor teams or on platform bring-up efforts
  • Exposure to model serving or inference optimization workflows

Benefits

  • Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities.
  • Competitive compensation packages, including RSU grants.
  • Annual target bonus.
  • Equity, with equity making up a significant portion of total compensation.
  • Relocation assistance provided for out-of-state candidates based in the US.
  • Earlier career stages work in a hybrid capacity at the Los Altos, CA office, with a minimum of 3 days per week on-site.
  • Senior members have both in office or remote flexibility.
  • Regular team onsites and local meetups in Los Altos, CA as well as different cities.

More jobs like this