CoreWeave logo
CoreWeave
Posted 176 days agoVerified live 2d ago

Operations Engineer, Fleet Reliability

Brief overview

New York, NYIn-person
UndergradOr in progress
$83k–$110k/yrStated range
2+ yrsMinimum
42 H-1B approvalsDept. of Labor
Linux system administrationHardware and software troubleshootingSystem maintenanceScripting (bash, python, powershell)GrafanaPrometheusPromsql queriesObservability platformsData center operationsServer rack managementHVAC systemsFiber traysKubernetes administrationHPC administrationGPU workload management

About the company

CoreWeave logo
CoreWeavecoreweave.com

CoreWeave is a cloud-based AI infrastructure company offering GPU cloud services to simplify AI and machine learning workloads.

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
42H-1B approved
100%approval rate
8new H-1B hires
$231,000median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20232
20246
202522
202612
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20232
20241
202519
202617
Top sponsored roles
Senior Business Systems EngineerStaff Software EngineerSENIOR NETWORK ENGINEERHPC Network EngineerSTAFF PRODUCT MANAGER, AI PERFORMANCE AND BENCHMARKING

Job description

Summary

CoreWeave is The Essential Cloud for AI™, providing a platform for innovators to build and scale AI with confidence. The Operations Engineer, Fleet Reliability will be responsible for managing the uptime and provisioning of CoreWeave’s fleet of server nodes, troubleshooting issues, and maintaining high-performance computing clusters.

Responsibilities

  • Configure and maintain large-scale high-performance supercomputing clusters running state-of-the-art GPUs
  • Troubleshoot hardware and software issues; escalate and coordinate as needed with data center, network, hardware and platform teams to drive resolution
  • Monitor and analyze system performance and take appropriate remediation actions for cloud health
  • Approach your work with flexibility and optimism anticipating shifting business and technical priorities
  • Create and maintain documentation of team processes, knowledge and best practices for system management
  • Think critically about your day-to-day work and work collaboratively to improve team processes and efficiency
  • Participate in oncall rotations which include after hours and weekend work

Skills

  • Strong understanding of Linux system administration and internals
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably
  • Software development or scripting languages (bash, python, powershell, etc)
  • 2 + years of experience troubleshooting or administering data center or on-prem infrastructure (servers, storage, network or a mix)
  • Grafana, Prometheus, promsql queries or similar observability platforms
  • Data center environments including server racks, HVAC systems, fiber trays
  • Kubernetes administration
  • HPC - administering GPU-related workloads
  • Bachelor's degree in a related field or equivalent experience

Qualifications

Must Haves

  • Strong understanding of Linux system administration and internals
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably
  • Software development or scripting languages (bash, python, powershell, etc)

Nice to Haves

  • 2 + years of experience troubleshooting or administering data center or on-prem infrastructure (servers, storage, network or a mix)
  • Grafana, Prometheus, promsql queries or similar observability platforms
  • Data center environments including server racks, HVAC systems, fiber trays
  • Kubernetes administration
  • HPC - administering GPU-related workloads
  • Bachelor's degree in a related field or equivalent experience

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

More jobs like this