Cerebras logo
Cerebras
Posted 19 days agoVerified live 15h ago

Software Engineer, Kernel Reliability

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
C/C++PythonOperating SystemsComputer ArchitectureSystems ProgrammingDebuggingFailure AnalysisParallel and Distributed ProgrammingDebugging Distributed and Parallel ApplicationsDebugging and Diagnostic ToolsMonitoringIncident ResponseRoot-Cause Analysis

Job description

Summary

Cerebras Systems builds advanced AI chips and computing systems designed to accelerate training and inference. The Software Engineer, Kernel Reliability will improve the reliability of compute clusters and underlying inference, training, and production services through systems programming, debugging, failure analysis, tooling, incident response, and collaboration across software, operations, ASIC, and hardware architecture teams.

Responsibilities

  • Contribute to the technical roadmap and execution for kernel-centric reliability of our internal and customer-facing systems
  • Partner with System and Cluster Operations teams to reduce system and service downtime after failure through tooling, analysis, and hands-on debugging support
  • Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis
  • Collaborate with software teams to improve the software stack—including kernels—to improve on-field debugging and failure analysis
  • Work with ASIC and hardware architecture teams to co-design next-generation architectures with reliability and ease of debug in mind
  • Participate in incident response, root-cause analysis, and post-mortems; drive follow-ups that measurably improve reliability over time

Skills

  • Strong programming skills in C/C++ and Python
  • Solid foundations in operating systems, computer architecture, and systems programming fundamentals
  • Ability to debug complex issues using logs, traces, and standard debugging workflows; interest in root-cause analysis
  • Exposure to parallel and distributed programming (message passing, multicore, GPU, embedded, etc.)
  • Experience building or using debug/diagnostic tools (debuggers, core dump handling, tracing, sanitizers, profilers, etc.)
  • Familiarity with debugging distributed and parallel applications (deadlocks, livelocks, race conditions, etc.)
  • Knowledge of computer architecture concepts (instruction pipelining, multithreading, networking, memory systems, etc.)
  • Operations & Monitoring: familiarity with monitoring, incident response, and post-mortem culture

Qualifications

Must Haves

  • Strong programming skills in C/C++ and Python
  • Solid foundations in operating systems, computer architecture, and systems programming fundamentals
  • Ability to debug complex issues using logs, traces, and standard debugging workflows; interest in root-cause analysis

Nice to Haves

  • Exposure to parallel and distributed programming (message passing, multicore, GPU, embedded, etc.)
  • Experience building or using debug/diagnostic tools (debuggers, core dump handling, tracing, sanitizers, profilers, etc.)
  • Familiarity with debugging distributed and parallel applications (deadlocks, livelocks, race conditions, etc.)
  • Knowledge of computer architecture concepts (instruction pipelining, multithreading, networking, memory systems, etc.)
  • Operations & Monitoring: familiarity with monitoring, incident response, and post-mortem culture

Benefits

  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • A simple, non-corporate work culture that respects individual beliefs.
  • A work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

More jobs like this