Summary
Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems. They are seeking a Network Operations Center (NOC) Operator to join their overnight operations team, focusing on monitoring, alert triage, and escalation to ensure reliable operations across their infrastructure.
Responsibilities
- Monitor data center systems using dashboards and alerting tools
- Acknowledge and triage alerts across compute, network, and hardware systems
- Identify and filter out false positives using predefined guidelines
- Follow structured runbooks to perform initial validation steps (e.g., running diagnostic scripts or checking system status)
- Escalate issues to the appropriate teams (hardware, network, SRE) based on clear escalation criteria
- Notify relevant stakeholders during active incidents
- Create, update, and manage tickets with accurate and timely information
- Attach relevant logs, diagnostics, and observations to support faster resolution
- Track incidents to ensure proper ownership and handoff
- Identify recurring alerts or patterns and report them to improve monitoring and reliability
- Maintain awareness of ongoing incidents and system status during your shift
- Gain exposure to large-scale AI and HPC infrastructure, including GPU-based systems
- Learn the fundamentals of data center operations, networking, and system reliability
- Opportunities to assist with hands-on data center tasks such as hardware checks, inspections, and basic support activities under guidance
Skills
- Basic familiarity with computers, Linux systems, or IT environments (academic or personal experience is acceptable)
- Ability to follow structured procedures and runbooks with attention to detail
- Strong communication skills and ability to clearly document issues
- Ability to prioritize tasks and respond quickly in a fast-paced, 24/7 environment
- Willingness to learn and grow in a technical operations role
- Exposure to monitoring tools (Grafana, Datadog, etc.)
- Basic understanding of networking or server hardware concepts
- Interest in data centers, cloud infrastructure, or AI systems
- Strong sense of ownership and urgency, with the confidence to escalate issues to the appropriate teams or levels to ensure timely resolution
Qualifications
Must Haves
- Basic familiarity with computers, Linux systems, or IT environments (academic or personal experience is acceptable)
- Ability to follow structured procedures and runbooks with attention to detail
- Strong communication skills and ability to clearly document issues
- Ability to prioritize tasks and respond quickly in a fast-paced, 24/7 environment
- Willingness to learn and grow in a technical operations role
Nice to Haves
- Exposure to monitoring tools (Grafana, Datadog, etc.)
- Basic understanding of networking or server hardware concepts
- Interest in data centers, cloud infrastructure, or AI systems
- Strong sense of ownership and urgency, with the confidence to escalate issues to the appropriate teams or levels to ensure timely resolution