Summary
Cornelis Networks develops high-performance scale-out networking solutions for AI and HPC datacenters, integrating hardware, software, and system-level technologies for compute clusters. The Software Engineer will develop and optimize AI/HPC communication middleware, integrate it with transport layers, contribute to open-source projects, collaborate across the software and hardware stack, and help diagnose customer performance issues.
Responsibilities
- AI/HPC Middleware Enablement & Optimization: Implement and help optimize features for HPC middleware (HPI and SHMEM) and AI middleware (CCL stacks ie: NCCL/RCCL and related collective communication libraries)
- Transport & Integration: Assist in integrating middleware capabilities with underlying transports and provider layers (ie: libfabric / OFI, UCX, verbs-like semantics where applicable)
- Upstream & Open-Source Leadership: Prepare and submit upstream contributions across MPI/SHMEM projects, CCL ecosystems, and related components
- Cross-Stack Collaboration & Performance Validation: Collaborate with kernel/driver and switch teams to help deliver end-to-end performance aligned to the Cornelis product roadmap
- Customer & Field Impact: Help analyze performance traces and reproduce customer issues, and assist in translating findings into fixes
Skills
- * 2+ years of experience in systems programming in C/C++ on Linux; Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)
- * Exposure to or coursework/project experience with HPC middleware (e.g., MPI/SHMEM) and/or AI collective communication libraries (e.g., NCCL/RCCL, CUDA/ROCm) and/or other CCL stacks
- * Familiarity with performance concepts and an ability to diagnose issues using profiling/tracing tools, or a strong willingness to learn
- * Foundational understanding of networking and/or RDMA concepts
- Location: This is a remote position for employees residing within the United States
- * Experience contributing to open-source projects
- * Exposure to libfabric providers
- * Familiarity with Ultra Ethernet (UEC/UET) specifications
- * Experience with RoCEv2, congestion control, and/or Ethernet-based RDMA deployments
- * Experience with benchmarking, profiling, and optimization
- * Background with Omni-Path/OPX or Ethernet-based HPC fabrics
Qualifications
Must Haves
- * 2+ years of experience in systems programming in C/C++ on Linux; Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience)
- * Exposure to or coursework/project experience with HPC middleware (e.g., MPI/SHMEM) and/or AI collective communication libraries (e.g., NCCL/RCCL, CUDA/ROCm) and/or other CCL stacks
- * Familiarity with performance concepts and an ability to diagnose issues using profiling/tracing tools, or a strong willingness to learn
- * Foundational understanding of networking and/or RDMA concepts
- Location: This is a remote position for employees residing within the United States
Nice to Haves
- * Experience contributing to open-source projects
- * Exposure to libfabric providers
- * Familiarity with Ultra Ethernet (UEC/UET) specifications
- * Experience with RoCEv2, congestion control, and/or Ethernet-based RDMA deployments
- * Experience with benchmarking, profiling, and optimization
- * Background with Omni-Path/OPX or Ethernet-based HPC fabrics
Benefits
- Equity
- Performance-based incentives, including an annual bonus or sales incentives
- Medical coverage
- Dental coverage
- Vision coverage
- Disability insurance
- Life insurance
- Dependent care flexible spending account
- Accidental injury insurance
- Pet insurance
- Generous paid holidays
- 401(k) with company match
- Open Time Off (OTO) for regular full-time exempt employees
- Sick time
- Bonding leave
- Pregnancy disability leave
- Remote work for employees residing within the United States
- Flexible work environment