Summary
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. They are seeking a Site Reliability Engineer to solve challenging problems in the Raptor engine organization, focusing on systems engineering issues to accelerate rocket engine development. The role involves managing server infrastructure, coordinating with IT teams, and supporting application deployment for optimal software performance.
Responsibilities
- Manage server infrastructure, HPC systems, storage systems, networks, and high-speed interconnect (Infiniband)
- Design, procure, and integrate infrastructure systems
- Work with propulsion engineering staff to solve critical bottlenecks
- Coordinate and communicate with company -wide infrastructure and IT teams
- Support application deployment (ANSYS, StarCCM+, manufacturing software) for best performance of software on real-world systems
Skills
- 1+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
- Bachelor's degree in computer science, engineering, math, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree
- Experience with Linux and Windows server software
- 1+ year of systems engineering experience
- Experience with scripting languages (Bash, Python), automation (Puppet, Ansible), and other common sysadmin tools
- Experience building, deploying, and troubleshooting large-scale compute systems
- Familiarity with resource development and management (Kubernetes, Docker)
- Familiarity with engineering and analysis applications, such as CFD and FEA
- Familiarity with diagnosing bottlenecks and designing systems for performance
- Able to work effectively in a dynamic environment while assuming high levels of responsibility and demonstrating accountability for rocket engine-level outcomes
Qualifications
Must Haves
- 1+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
- Bachelor's degree in computer science, engineering, math, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree
- Experience with Linux and Windows server software
Nice to Haves
- 1+ year of systems engineering experience
- Experience with scripting languages (Bash, Python), automation (Puppet, Ansible), and other common sysadmin tools
- Experience building, deploying, and troubleshooting large-scale compute systems
- Familiarity with resource development and management (Kubernetes, Docker)
- Familiarity with engineering and analysis applications, such as CFD and FEA
- Familiarity with diagnosing bottlenecks and designing systems for performance
- Able to work effectively in a dynamic environment while assuming high levels of responsibility and demonstrating accountability for rocket engine-level outcomes
Benefits
- Long-term incentives, in the form of company stock, stock options, or long-term cash awards
- Potential discretionary bonuses
- Ability to purchase additional stock at a discount through an Employee Stock Purchase Plan
- Access to comprehensive medical, vision, and dental coverage
- Access to a 401(k) retirement plan
- Short and long-term disability insurance
- Life insurance
- Paid parental leave
- Various other discounts and perks
- 3 weeks of paid vacation
- Eligible for 10 or more paid holidays per year
- Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law