Summary
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. The role involves developing and maintaining production systems, automating infrastructure, and collaborating across teams to enhance machine lifecycle processes.
Responsibilities
- Develop and Maintain Production Systems: Design, implement, and improve software that powers GPU fleet lifecycle management and machine configuration at scale
- Automate Infrastructure: Build and enhance automation frameworks for machine provisioning, configuration management, and deployment
- Support New Hardware Introduction (NPI): Enable bring-up, validation, and production readiness for new server and accelerator platforms
- Enhance Machine Lifecycle Processes: Improve and refine workflows for bare metal provisioning, firmware updates, and system health monitoring
- Debug Hardware and Firmware Issues: Investigate failures across BIOS, BMC, firmware, networking, storage, and boot flows
- Collaborate Across Teams: Work closely with infrastructure, security, and product engineering teams to develop scalable and maintainable solutions
Skills
- Have 2+ years of experience working with Go (Golang) or Python in production environments
- Have 2+ years of experience with configuration management tools and practices
- Are comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers
- Can independently troubleshoot complex systems and communicate effectively across software, infrastructure, and vendor teams
- Experience with Go in infrastructure, systems, or backend development
- Hands-on experience with bare metal provisioning and lifecycle management, including technologies such as Redfish, BMC, IPMI, DHCP, and PXE
- Experience diagnosing issues involving drivers, firmware, and hardware compatibility across GPU servers
- Experience incorporating AI-assisted development tools into engineering workflows, including code generation, debugging, test development, and documentation
- Experience building Linux distributions or managing OS customization and imaging
- Familiarity with Ansible for system configuration and automation
- Exposure to Kubernetes and container orchestration concepts
Qualifications
Must Haves
- Have 2+ years of experience working with Go (Golang) or Python in production environments
- Have 2+ years of experience with configuration management tools and practices
- Are comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers
- Can independently troubleshoot complex systems and communicate effectively across software, infrastructure, and vendor teams
Nice to Haves
- Experience with Go in infrastructure, systems, or backend development
- Hands-on experience with bare metal provisioning and lifecycle management, including technologies such as Redfish, BMC, IPMI, DHCP, and PXE
- Experience diagnosing issues involving drivers, firmware, and hardware compatibility across GPU servers
- Experience incorporating AI-assisted development tools into engineering workflows, including code generation, debugging, test development, and documentation
- Experience building Linux distributions or managing OS customization and imaging
- Familiarity with Ansible for system configuration and automation
- Exposure to Kubernetes and container orchestration concepts
Benefits
- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use