Summary
Cerebras Systems builds the world's largest AI chip, transforming the user experience of AI applications. As a Kernel Engineer, you will develop high-performance software for cutting-edge AI and high-performance computing workloads, collaborating with various engineering teams to optimize and validate machine learning and linear algebra operations.
Responsibilities
- Help design and implement machine learning and linear algebra kernels for the Cerebras Wafer-Scale Engine
- Develop and debug high-performance kernel routines using low-level programming techniques and the Cerebras Software Language, a custom C-like language
- Apply parallel programming algorithms to map computational workloads efficiently onto the Cerebras architecture
- Use mathematical analysis, performance data, and profiling tools to evaluate kernel behavior and inform design decisions
- Identify and investigate correctness, performance, and hardware utilization issues
- Develop unit tests and system-level validation methodologies to verify the functionality and performance of kernel libraries
- Collaborate with kernel, compiler, performance, and hardware engineers to improve software and system performance
- Study emerging machine learning workloads and contribute to the evolution of the kernel library
- Participate in code reviews, technical discussions, and software development processes
- Build an understanding of the Cerebras architecture, instruction set, memory system, and communication model
Skills
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field
- Strong programming fundamentals in C++ and familiarity with Python
- Understanding of foundational computer architecture concepts such as processors, memory hierarchies, instruction execution, or data movement
- Knowledge of data structures, algorithms, and software development fundamentals
- Experience debugging software through coursework, internships, research, co-op placements, or technical projects
- Strong analytical and problem-solving skills
- Interest in low-level software, parallel computing, performance optimization, or hardware/software co-design
- Ability to learn unfamiliar systems and collaborate effectively within a technical team
- Research, internships, or projects involving kernel development, compilers, computer architecture, HPC, or systems programming
- Familiarity with parallel algorithms, multithreaded programming, or distributed memory systems
- Exposure to programming accelerators such as GPUs, FPGAs, or other specialized processors
- Experience with low-level programming, assembly language, CUDA, OpenCL, or a domain-specific language
- Familiarity with machine learning concepts, neural networks, or frameworks such as PyTorch or TensorFlow
- Exposure to numerical computing, linear algebra, or HPC kernels
- Experience using profiling, benchmarking, or performance analysis tools
- Familiarity with library or API development practices
Qualifications
Must Haves
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field
- Strong programming fundamentals in C++ and familiarity with Python
- Understanding of foundational computer architecture concepts such as processors, memory hierarchies, instruction execution, or data movement
- Knowledge of data structures, algorithms, and software development fundamentals
- Experience debugging software through coursework, internships, research, co-op placements, or technical projects
- Strong analytical and problem-solving skills
- Interest in low-level software, parallel computing, performance optimization, or hardware/software co-design
- Ability to learn unfamiliar systems and collaborate effectively within a technical team
Nice to Haves
- Research, internships, or projects involving kernel development, compilers, computer architecture, HPC, or systems programming
- Familiarity with parallel algorithms, multithreaded programming, or distributed memory systems
- Exposure to programming accelerators such as GPUs, FPGAs, or other specialized processors
- Experience with low-level programming, assembly language, CUDA, OpenCL, or a domain-specific language
- Familiarity with machine learning concepts, neural networks, or frameworks such as PyTorch or TensorFlow
- Exposure to numerical computing, linear algebra, or HPC kernels
- Experience using profiling, benchmarking, or performance analysis tools
- Familiarity with library or API development practices