Summary
Prime Intellect is building an open superintelligence stack that provides infrastructure for frontier AI labs and teams. The Member of Technical Staff - Datacenter Operations will own the operational readiness of physical GPU-cloud infrastructure by coordinating deployments, maintenance, incident response, capacity readiness, and datacenter partners.
Responsibilities
- Coordinate rack deployment, cabling, inventory, and acceptance testing for new GPU capacity with datacenter partners and engineering teams
- Maintain accurate asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories
- Lead hardware fault triage and coordinate remote hands, vendor escalations, component replacement, and RMA workflows
- Establish maintenance plans and change procedures that minimize customer disruption and protect equipment and data
- Track capacity readiness, hardware failure trends, repair times, and operational risks; automate repetitive reporting and workflows
- Partner with facility teams on power, cooling, environmental monitoring, and readiness for high-density GPU deployments
- Create runbooks and escalation procedures and support incident response across datacenter and infrastructure teams
Skills
- 3+ years in datacenter operations, hardware infrastructure, or production systems operations
- Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling
- Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents
- Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools
- Strong operational judgment, documentation habits, and ownership of issues through resolution
- GPU server components, PCIe devices, memory, storage, and hardware diagnostics
- Rack power budgeting, redundant power paths, airflow, and high-density cooling fundamentals
- Fiber and copper cabling, optics, labeling, and physical network troubleshooting
- Asset tracking, spares management, change control, and incident management
- Basic scripting for inventory, health checks, and operational automation; familiarity with safe datacenter working practices
- Experience with NVIDIA DGX/HGX systems or large GPU cluster deployments
- Liquid-cooled infrastructure and coordination with facility engineering teams
- Experience bringing up new datacenter sites or expanding multi-site capacity
- Hardware qualification, burn-in testing, and reliability analysis
- Experience integrating physical operations with automated fleet provisioning
Qualifications
Must Haves
- 3+ years in datacenter operations, hardware infrastructure, or production systems operations
- Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling
- Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents
- Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools
- Strong operational judgment, documentation habits, and ownership of issues through resolution
- GPU server components, PCIe devices, memory, storage, and hardware diagnostics
- Rack power budgeting, redundant power paths, airflow, and high-density cooling fundamentals
- Fiber and copper cabling, optics, labeling, and physical network troubleshooting
- Asset tracking, spares management, change control, and incident management
- Basic scripting for inventory, health checks, and operational automation; familiarity with safe datacenter working practices
Nice to Haves
- Experience with NVIDIA DGX/HGX systems or large GPU cluster deployments
- Liquid-cooled infrastructure and coordination with facility engineering teams
- Experience bringing up new datacenter sites or expanding multi-site capacity
- Hardware qualification, burn-in testing, and reliability analysis
- Experience integrating physical operations with automated fleet provisioning
Benefits