Diverse Lynx logo
Diverse Lynx
Posted 10 days agoVerified live 1d ago

HPC Consultant

Brief overview

Remote
UndergradOr in progress
$120k/yrStated minimum
4+ yrsMinimum
1 H-1B approvalsDept. of Labor
KubernetesTerraformAWSPythonHigh-Performance Computing (HPC)GPU Computing

About the company

Diverse Lynx logo
Diverse Lynxdiverselynx.com

Diverse Lynx is a WBENC- and NMSDC-certified partner, helping organizations turn diversity goals into measurable impact through staffing and contingent workforce solutions.

Visa sponsorship history

1 year sponsoring, last filed FY2025

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
1H-1B approved
100%approval rate
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20251

Job description

Summary

Confidential is seeking an HPC Consultant to operate large-scale, multi-cloud Kubernetes and HPC infrastructure. The role focuses on cluster lifecycle management, GPU job scheduling, infrastructure provisioning, monitoring, incident response, and collaboration with networking, storage, security, and AI/ML teams.

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth
  • Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Eligible to workP, and OCI, with additional providers to be expanded in the near future
  • Manage job scheduling to allocate GPU compute across training and inference workloads
  • Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews
  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams

Skills

  • · 4+ years in infrastructure engineering, cloud platforms, or HPC
  • · **Kubernetes is the core requirement.** You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit
  • · Terraform proficiency. You'll write and review infrastructure-as-code daily
  • · Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
  • · Python for tooling and automation

Qualifications

Must Haves

  • · 4+ years in infrastructure engineering, cloud platforms, or HPC
  • · **Kubernetes is the core requirement.** You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit
  • · Terraform proficiency. You'll write and review infrastructure-as-code daily
  • · Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
  • · Python for tooling and automation

Benefits

  • Remote-first work model with occasional on-site visits to customer offices
  • Benefits

More jobs like this