Orion Placement logo
Orion Placement
Posted 37 days agoVerified live 9h ago

GPU Infrastructure NOC Engineer

Brief overview

Remote
$75k–$140k/yrStated range
2+ yrsMinimum
NOC OperationsInfrastructure MonitoringGPU InfrastructureHigh-Performance Computing (HPC)PythonBashDatadogGrafanaPagerDutyIncident ResponseNetwork InfrastructureData Center Infrastructure

About the company

Orion Placement logo
Orion Placementorionplacement.com

Orion Placement provides recruitment and staffing services for professionals and support staff across multiple industries.

Job description

Summary

Orion Placement is hiring a GPU Infrastructure NOC Engineer to support next-generation AI infrastructure and high-performance GPU compute environments. The role monitors production GPU clusters, troubleshoots incidents, maintains SLA commitments, develops automation and monitoring tools, and coordinates with technical partners during customer-impacting issues.

Responsibilities

  • Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments
  • Triage, troubleshoot, and resolve incidents while maintaining SLA requirements
  • Identify infrastructure issues early and take proactive action before they become customer-impacting incidents
  • Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners
  • Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work
  • Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows
  • Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency
  • Create, maintain, and continuously improve technical runbooks and standard operating procedures
  • Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements
  • Track SLA and incident metrics and identify opportunities to improve reliability and response times
  • Communicate clearly and proactively with customers and internal stakeholders during incidents
  • Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues
  • Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts

Skills

  • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment
  • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments
  • Hands-on Python or Bash scripting experience
  • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms
  • Strong troubleshooting and incident-response skills
  • Experience working with network, compute, storage, or data center infrastructure
  • Ability to understand technical issues quickly and communicate effectively during incidents
  • Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement
  • Must be comfortable working rotating 24/7 shifts, including nights and weekends
  • Ability to work independently in a remote environment while collaborating effectively with global technical teams
  • Additional languages beyond English are a plus

Qualifications

Must Haves

  • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment
  • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments
  • Hands-on Python or Bash scripting experience
  • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms
  • Strong troubleshooting and incident-response skills
  • Experience working with network, compute, storage, or data center infrastructure
  • Ability to understand technical issues quickly and communicate effectively during incidents
  • Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement
  • Must be comfortable working rotating 24/7 shifts, including nights and weekends
  • Ability to work independently in a remote environment while collaborating effectively with global technical teams

Nice to Haves

  • Additional languages beyond English are a plus

Benefits

  • Bonus and equity opportunities in addition to competitive base compensation.
  • Fully remote role supporting a 24/7 global operations environment.
  • Remote nationwide flexibility with multiple shift options.
  • Medical insurance.
  • Dental insurance.
  • Vision insurance.
  • 401(k).
  • Paid maternity and paternity leave.
  • Paid time off.
  • Retirement plan.

More jobs like this