Summary
Bitdeer is a technology company building AI computational infrastructure and Bitcoin mining solutions. The GPU Compute & Bare Metal / DPU Engineer will own the end-to-end lifecycle of bare-metal GPU nodes across multiple regions, including provisioning, firmware and DPU management, break-fix, reliability, and decommissioning. The role will also automate node delivery, improve fleet operations, and collaborate with storage, imaging, and networking teams.
Responsibilities
- Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery / onboarding, in-service operation, break-fix, and decommissioning across multiple regions
- Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs
- Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation
- Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring
- Participate in a sustainable 7×24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil
- Partner with Storage / Image and Network teams to streamline the provisioning-to-handoff path
- Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions
Skills
- 3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations, HPC, or cloud infrastructure
- Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX-class), including driver / CUDA and firmware management
- Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning
- Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking
- Infrastructure automation skills (Ansible, Terraform, Python / Go)
- Comfortable owning on-call, incident management, and operational runbooks
- Multi-region / large-fleet operations experience is a strong plus
Qualifications
Must Haves
- 3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations, HPC, or cloud infrastructure
- Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX-class), including driver / CUDA and firmware management
- Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning
- Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking
- Infrastructure automation skills (Ansible, Terraform, Python / Go)
- Comfortable owning on-call, incident management, and operational runbooks
Nice to Haves
- Multi-region / large-fleet operations experience is a strong plus
Benefits
- Remote (within locations)