Summary
Prime Intellect is building an open superintelligence infrastructure stack for frontier AI research and production workloads. The Member of Technical Staff will build and operate reliable, high-throughput storage systems for datasets, checkpoints, and model artifacts, while managing performance, durability, availability, security, and cost at GPU-cluster scale.
Responsibilities
- Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows
- Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads
- Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads
- Build provisioning, capacity planning, lifecycle management, and operational automation for storage services
- Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets
- Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices
- Implement access controls, tenant separation, quotas, monitoring, and runbooks; collaborate with compute and networking teams
Skills
- • 3+ years building or operating production distributed storage systems
- • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS
- • Strong Linux administration and performance troubleshooting skills
- • Experience automating infrastructure operations in Python, Go, Bash, or similar languages
- • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
- • Block, file, and object storage semantics and their performance tradeoffs
- • NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
- • High-throughput storage networking and distributed client behavior
- • Capacity forecasting, observability, alerting, and safe maintenance procedures
- • Authentication, authorization, encryption, and secure data lifecycle management
- • Experience supporting large GPU training clusters and high-volume checkpoint workloads
- • S3-compatible object storage, data tiering, or distributed caching
- • RDMA-enabled storage or GPUDirect Storage experience
- • Kubernetes storage integrations or SLURM environments
- • Storage cost optimization and contributions to open-source storage systems
Qualifications
Must Haves
- • 3+ years building or operating production distributed storage systems
- • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS
- • Strong Linux administration and performance troubleshooting skills
- • Experience automating infrastructure operations in Python, Go, Bash, or similar languages
- • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
- • Block, file, and object storage semantics and their performance tradeoffs
- • NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
- • High-throughput storage networking and distributed client behavior
- • Capacity forecasting, observability, alerting, and safe maintenance procedures
- • Authentication, authorization, encryption, and secure data lifecycle management
Nice to Haves
- • Experience supporting large GPU training clusters and high-volume checkpoint workloads
- • S3-compatible object storage, data tiering, or distributed caching
- • RDMA-enabled storage or GPUDirect Storage experience
- • Kubernetes storage integrations or SLURM environments
- • Storage cost optimization and contributions to open-source storage systems
Benefits