Summary
MaintainX is a mobile-first work execution platform for industrial and frontline teams. The company is seeking a Site Reliability Engineer to improve platform reliability, observability, resilience, and developer autonomy by partnering with development teams, establishing reliability standards, and building shared tooling and practices.
Responsibilities
- Assess service maturity and provide insights to development teams
- Partner with development teams to implement observability best practices
- Enable development teams to become autonomous with their service deployment, support, and infrastructure
- Mentor developers on reliability practices, focusing on making them self-sufficient
- Act as the bridge, ear and eyes of the Platform Division teams to drive tooling and practice adoption across development teams
Skills
- Deep understanding of observability practices in a distributed system environment and how it influences system design and team behaviour
- Practical experience with SRE concepts (SLOs, error budgets, incident management)
- 3–5+ years in software development, SRE, DevOps, or production development roles with experience operating production systems
- Proficient in cloud-native platforms and infrastructure-as-code concepts and tools
- Working knowledge of at least one programming language (TypeScript/Node.js is a plus)
- Excellent communication and collaboration abilities across technical and non-technical teams
- Ability to translate complex reliability concepts into actionable guidance
- You enjoy enabling teams to succeed independently and measuring success by reduced dependency on you
Qualifications
Must Haves
- Deep understanding of observability practices in a distributed system environment and how it influences system design and team behaviour
- Practical experience with SRE concepts (SLOs, error budgets, incident management)
- 3–5+ years in software development, SRE, DevOps, or production development roles with experience operating production systems
- Proficient in cloud-native platforms and infrastructure-as-code concepts and tools
- Working knowledge of at least one programming language (TypeScript/Node.js is a plus)
- Excellent communication and collaboration abilities across technical and non-technical teams
- Ability to translate complex reliability concepts into actionable guidance
- You enjoy enabling teams to succeed independently and measuring success by reduced dependency on you
Benefits
- Depending on the role, compensation may also include commission, an annual bonus and equity.