Summary
Confidential is seeking an HPC Consultant to operate large-scale, multi-cloud Kubernetes and HPC infrastructure. The role focuses on cluster lifecycle management, GPU job scheduling, infrastructure provisioning, monitoring, incident response, and collaboration with networking, storage, security, and AI/ML teams.
Responsibilities
- Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth
- Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Eligible to workP, and OCI, with additional providers to be expanded in the near future
- Manage job scheduling to allocate GPU compute across training and inference workloads
- Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews
- Coordinate daily with Networking, Storage, Security, and AI/ML platform teams
Skills
- · 4+ years in infrastructure engineering, cloud platforms, or HPC
- · **Kubernetes is the core requirement.** You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit
- · Terraform proficiency. You'll write and review infrastructure-as-code daily
- · Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
- · Python for tooling and automation
Qualifications
Must Haves
- · 4+ years in infrastructure engineering, cloud platforms, or HPC
- · **Kubernetes is the core requirement.** You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit
- · Terraform proficiency. You'll write and review infrastructure-as-code daily
- · Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre)
- · Python for tooling and automation
Benefits
- Remote-first work model with occasional on-site visits to customer offices
- Benefits