Summary
Hyperbolic Labs is building an Open-Access AI Cloud and GPU marketplace to make computing power more affordable and accessible. The Technical Support Engineer owns customer support tickets end to end, troubleshoots Linux-based GPU infrastructure, manages SLA commitments, and develops runbooks and documentation while coordinating escalations when needed.
Responsibilities
- Ticket ownership, end to end. You own every ticket you pick up, including after it escalates. You do not hand off, you pull in the engineer you need and stay on it until the customer is working
- The SLA clock. First response, severity classification, and keeping us honest against our response commitments. You are the person who knows where every open issue stands
- First technical response and triage. Reproduce the problem, gather the logs and configuration that matter, and make the first call on whether the fault is ours or the provider's
- Customer environment access and configuration. SSH key and access issues, NFS mounts and storage, quotas, security groups, container and driver questions, billing and account questions
- Runbook execution and authoring. Run the documented play when there is one. When you solve something new, write the runbook so the next person does not escalate it
- Documentation. Keep our customer-facing docs and internal knowledge base current. Most repeat tickets are a documentation gap
Skills
- Very strong Linux experience and daily work in the CLI
- Experience owning tickets against a response SLA in cloud, hosting, or infrastructure support
- Solid networking and storage fundamentals: SSH, NFS and mounts, DNS, firewalls and security groups
- Working familiarity with GPU workloads: nvidia-smi, drivers, CUDA, containers
- Clear and fast written communication under time pressure. Customers read what you write while they are blocked
- Good judgment about the limits of your own knowledge, and a bias toward escalating early with a complete picture rather than late with a guess
- Comfortable working across time zones and with an on-call rotation for critical issues
- Experience with ticketing and on-call tooling (Zendesk, Linear, PagerDuty, or similar)
- Scripting in bash or Python to automate repeat work
- Exposure to Slurm, Kubernetes, or Docker in a multi-tenant environment
- Background in GPU cloud, HPC, or a hardware-adjacent support org
Qualifications
Must Haves
- Very strong Linux experience and daily work in the CLI
- Experience owning tickets against a response SLA in cloud, hosting, or infrastructure support
- Solid networking and storage fundamentals: SSH, NFS and mounts, DNS, firewalls and security groups
- Working familiarity with GPU workloads: nvidia-smi, drivers, CUDA, containers
- Clear and fast written communication under time pressure. Customers read what you write while they are blocked
- Good judgment about the limits of your own knowledge, and a bias toward escalating early with a complete picture rather than late with a guess
- Comfortable working across time zones and with an on-call rotation for critical issues
Nice to Haves
- Experience with ticketing and on-call tooling (Zendesk, Linear, PagerDuty, or similar)
- Scripting in bash or Python to automate repeat work
- Exposure to Slurm, Kubernetes, or Docker in a multi-tenant environment
- Background in GPU cloud, HPC, or a hardware-adjacent support org
Benefits
- Employees who do this work well can move into the Forward Deployed Engineer role, with that career path deliberately built with them.