Summary
Prime Intellect is building an open superintelligence stack that provides infrastructure for frontier AI development. The Member of Technical Staff will design and operate high-performance datacenter networks connecting large GPU clusters, ensuring reliable and scalable training, storage, inference, and management connectivity.
Responsibilities
- Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic
- Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with clear standards for routing, redundancy, and capacity
- Automate network provisioning, configuration validation, upgrades, and rollback procedures
- Diagnose packet loss, congestion, link failures, and collective communication performance across hosts and switches
- Benchmark end-to-end network performance with infrastructure and ML teams, translating workload needs into measurable acceptance criteria
- Build monitoring for port health, errors, utilization, congestion, and fabric topology; improve incident response and runbooks
- Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution
Skills
- 3+ years of production datacenter networking experience
- Strong understanding of Ethernet, TCP/IP, routing, switching, and redundant network design
- Hands-on experience with high-performance GPU networking using InfiniBand or RoCE
- Experience troubleshooting network problems across Linux hosts, NICs, switches, and physical links
- Ability to automate network operations with Python, Ansible, or comparable tools
- Leaf-spine architectures, BGP, ECMP, VLANs, and network segmentation
- RDMA concepts and performance tuning; congestion control and lossless Ethernet considerations
- Linux networking, NIC drivers and firmware, packet capture, and throughput/latency testing
- Optics, transceivers, cable management, and link-level diagnostics
- Safe change management, configuration versioning, telemetry, and alerting
- Experience operating 400G/800G networks or large multi-rack GPU clusters
- NVIDIA Spectrum or Quantum networking experience
- NCCL performance analysis and distributed training troubleshooting
- EVPN/VXLAN, SONiC, or network source-of-truth systems
- Experience with network simulation, automated validation, and capacity planning
Qualifications
Must Haves
- 3+ years of production datacenter networking experience
- Strong understanding of Ethernet, TCP/IP, routing, switching, and redundant network design
- Hands-on experience with high-performance GPU networking using InfiniBand or RoCE
- Experience troubleshooting network problems across Linux hosts, NICs, switches, and physical links
- Ability to automate network operations with Python, Ansible, or comparable tools
- Leaf-spine architectures, BGP, ECMP, VLANs, and network segmentation
- RDMA concepts and performance tuning; congestion control and lossless Ethernet considerations
- Linux networking, NIC drivers and firmware, packet capture, and throughput/latency testing
- Optics, transceivers, cable management, and link-level diagnostics
- Safe change management, configuration versioning, telemetry, and alerting
Nice to Haves
- Experience operating 400G/800G networks or large multi-rack GPU clusters
- NVIDIA Spectrum or Quantum networking experience
- NCCL performance analysis and distributed training troubleshooting
- EVPN/VXLAN, SONiC, or network source-of-truth systems
- Experience with network simulation, automated validation, and capacity planning
Benefits