Bitdeer (NASDAQ: BTDR) logo
Bitdeer (NASDAQ: BTDR)
Posted 52 days agoVerified live 2d ago

SRE L1 Support/Cloud Platform Ops Engineer

Brief overview

Remote
2+ yrsMinimum
2 H-1B approvalsDept. of Labor
Linux System AdministrationLog AnalysisService ManagementPrometheusGrafanaNagiosServiceNowJira Service ManagementRack and StackData Center CablingHardware Replacement

About the company

Bitdeer (NASDAQ: BTDR) logo
Bitdeer (NASDAQ: BTDR)bitdeer.com

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Visa sponsorship history

3 years sponsoring, last filed FY2025

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
2H-1B approved
100%approval rate
1new H-1B hires
$132,000median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20231
20251
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20242
20252
Top sponsored roles
Legal CounselSenior Legal CounselDirector of Government Relations (North America Region)

Job description

Summary

Bitdeer is a technology company that provides Bitcoin mining solutions and builds AI computational infrastructure and cloud capabilities. The SRE L1 Support/Cloud Platform Ops Engineer will monitor and support US GPU data centers, respond to incidents, execute remediation runbooks, perform hardware triage, manage tickets, and provide structured operational data to improve AIOps automation.

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports
  • Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira)
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles)
  • Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team
  • Maintain and update operational runbooks based on recurring issues
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance
  • Every novel incident you resolve is data the platform team needs — you tag it, describe it, and hand it back so it becomes an automation
  • Every runbook you touch should get closer to being executable by the platform, not by you
  • Your handoff notes are structured signal, not free-form email

Skills

  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change
  • Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed

Qualifications

Must Haves

  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation — you don't just execute the runbook, you notice when it's the third time this month and ask what should change
  • Comfort with structured data — you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed

Benefits

  • Remote work within San Jose, CA, United States or Austin, TX, United States
  • Growth path into SME roles or the platform team as automation authors

More jobs like this