Dice logo
Dice
Posted 10 days agoVerified live 1d ago

HPC (High-Performance Computing) Consultant

Brief overview

Remote
UndergradOr in progress
4+ yrsMinimum
KubernetesTerraformAmazon Web ServicesAmazon EC2Amazon S3Amazon EFSAmazon FSx for LustrePythonHigh-Performance ComputingGPU ComputingContinuous Integration/Continuous DeliveryJob Scheduling

About the company

Dice is the go-to career marketplace for tech professionals.

Job description

Summary

Dice is recruiting for an HPC Consultant to support large-scale, multi-cloud computing infrastructure. The role focuses on operating Kubernetes platforms, provisioning HPC infrastructure, managing GPU job scheduling, maintaining observability and incident response processes, and coordinating with networking, storage, security, and AI/ML teams.

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth
  • Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be expanded in the near future
  • Manage job scheduling to allocate GPU compute across training and inference workloads
  • Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews
  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams

Skills

  • 4+ Years in infrastructure engineering, cloud platforms, or HPC
  • Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit
  • Terraform proficiency. You'll write and review infrastructure-as-code daily
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
  • Python for tooling and automation

Qualifications

Must Haves

  • 4+ Years in infrastructure engineering, cloud platforms, or HPC
  • Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit
  • Terraform proficiency. You'll write and review infrastructure-as-code daily
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
  • Python for tooling and automation

Benefits

  • Remote work arrangement (Remote USA)

More jobs like this