Prime Intellect logo
Prime Intellect
Posted 17 days agoVerified live 12h ago

Member of Technical Staff - Storage Infrastructure

Brief overview

Remote
UndergradOr in progress
$150k–$300k/yrStated range
3+ yrsMinimum
Distributed Storage SystemsLustre, BeeGFS, Ceph, or GPFSLinux AdministrationStorage Performance TroubleshootingPythonGoBashData Replication and RecoveryNVMe/SSD PerformanceFilesystem TuningI/O Profiling and BenchmarkingStorage Networking

About the company

Prime Intellect logo
Prime Intellectprimeintellect.ai

An organization building open-source AI infrastructure for decentralized, scalable ML systems.

Job description

Summary

Prime Intellect is building an open superintelligence infrastructure stack for frontier AI research and production workloads. The Member of Technical Staff will build and operate reliable, high-throughput storage systems for datasets, checkpoints, and model artifacts, while managing performance, durability, availability, security, and cost at GPU-cluster scale.

Responsibilities

  • Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows
  • Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads
  • Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads
  • Build provisioning, capacity planning, lifecycle management, and operational automation for storage services
  • Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks; collaborate with compute and networking teams

Skills

  • • 3+ years building or operating production distributed storage systems
  • • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS
  • • Strong Linux administration and performance troubleshooting skills
  • • Experience automating infrastructure operations in Python, Go, Bash, or similar languages
  • • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
  • • Block, file, and object storage semantics and their performance tradeoffs
  • • NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
  • • High-throughput storage networking and distributed client behavior
  • • Capacity forecasting, observability, alerting, and safe maintenance procedures
  • • Authentication, authorization, encryption, and secure data lifecycle management
  • • Experience supporting large GPU training clusters and high-volume checkpoint workloads
  • • S3-compatible object storage, data tiering, or distributed caching
  • • RDMA-enabled storage or GPUDirect Storage experience
  • • Kubernetes storage integrations or SLURM environments
  • • Storage cost optimization and contributions to open-source storage systems

Qualifications

Must Haves

  • • 3+ years building or operating production distributed storage systems
  • • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS
  • • Strong Linux administration and performance troubleshooting skills
  • • Experience automating infrastructure operations in Python, Go, Bash, or similar languages
  • • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
  • • Block, file, and object storage semantics and their performance tradeoffs
  • • NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
  • • High-throughput storage networking and distributed client behavior
  • • Capacity forecasting, observability, alerting, and safe maintenance procedures
  • • Authentication, authorization, encryption, and secure data lifecycle management

Nice to Haves

  • • Experience supporting large GPU training clusters and high-volume checkpoint workloads
  • • S3-compatible object storage, data tiering, or distributed caching
  • • RDMA-enabled storage or GPUDirect Storage experience
  • • Kubernetes storage integrations or SLURM environments
  • • Storage cost optimization and contributions to open-source storage systems

Benefits

  • Equity incentives

More jobs like this