Voltage Park logo
Voltage Park
Posted 138 days agoVerified live 16h ago

Network Operations Center (NOC) Operator

Brief overview

Allen, TXIn-person
$75k–$85k/yrStated range
6 H-1B approvalsDept. of Labor
Following structured procedures and runbooksStrong communication skillsClear documentation of issuesTask prioritizationQuick response in fast-paced environmentsIncident triage and escalationTicket managementLog and diagnostic analysisMonitoring tools (Grafana, Datadog, etc.)Basic networking conceptsServer hardware conceptsData center operationsSystem reliabilityHands-on hardware support

About the company

Voltage Park logo
Voltage Parkvoltagepark.com

Voltage Park provides infrastructure for machine learning.

Visa sponsorship history

2 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
6H-1B approved
100%approval rate
$190,000median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20255
20261
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20253
20261
Top sponsored roles
Software EngineerStaff Software EngineerSenior Software Engineer

Job description

Summary

Lightning AI is the company behind PyTorch Lightning, combining developer-first software with large-scale compute infrastructure through its merger with Voltage Park. They are seeking a Network Operations Center (NOC) Operator to monitor and respond to system alerts, ensuring reliable operations across their data center infrastructure.

Responsibilities

  • Monitor data center systems using dashboards and alerting tools
  • Acknowledge and triage alerts across compute, network, and hardware systems
  • Identify and filter out false positives using predefined guidelines
  • Follow structured runbooks to perform initial validation steps (e.g., running diagnostic scripts or checking system status)
  • Escalate issues to the appropriate teams (hardware, network, SRE) based on clear escalation criteria
  • Notify relevant stakeholders during active incidents
  • Create, update, and manage tickets with accurate and timely information
  • Attach relevant logs, diagnostics, and observations to support faster resolution
  • Track incidents to ensure proper ownership and handoff
  • Identify recurring alerts or patterns and report them to improve monitoring and reliability
  • Maintain awareness of ongoing incidents and system status during your shift
  • Gain exposure to large-scale AI and HPC infrastructure, including GPU-based systems
  • Learn the fundamentals of data center operations, networking, and system reliability
  • Opportunities to assist with hands-on data center tasks such as hardware checks, inspections, and basic support activities under guidance

Skills

  • Basic familiarity with computers, Linux systems, or IT environments (academic or personal experience is acceptable)
  • Ability to follow structured procedures and runbooks with attention to detail
  • Strong communication skills and ability to clearly document issues
  • Ability to prioritize tasks and respond quickly in a fast-paced, 24/7 environment
  • Willingness to learn and grow in a technical operations role
  • Exposure to monitoring tools (Grafana, Datadog, etc.)
  • Basic understanding of networking or server hardware concepts
  • Interest in data centers, cloud infrastructure, or AI systems
  • Strong sense of ownership and urgency, with the confidence to escalate issues to the appropriate teams or levels to ensure timely resolution

Qualifications

Must Haves

  • Basic familiarity with computers, Linux systems, or IT environments (academic or personal experience is acceptable)
  • Ability to follow structured procedures and runbooks with attention to detail
  • Strong communication skills and ability to clearly document issues
  • Ability to prioritize tasks and respond quickly in a fast-paced, 24/7 environment
  • Willingness to learn and grow in a technical operations role

Nice to Haves

  • Exposure to monitoring tools (Grafana, Datadog, etc.)
  • Basic understanding of networking or server hardware concepts
  • Interest in data centers, cloud infrastructure, or AI systems
  • Strong sense of ownership and urgency, with the confidence to escalate issues to the appropriate teams or levels to ensure timely resolution

More jobs like this