Summary
System Automation Corporation develops SaaS solutions for state and local regulatory agencies. The company is seeking a mid-level Site Reliability Engineer to build and operate critical Microsoft Azure cloud infrastructure, establish observability standards, automate operational processes, maintain CI/CD pipelines, and improve platform reliability through incident response and postmortems.
Responsibilities
- Build, operate, and scale production systems in Microsoft Azure (App Service, Networking, WAF, CosmosDB and related infrastructure) to meet availability and performance targets
- Participate in an on-call rotation; triage, respond to, and resolve production incidents, and lead or contribute to blameless postmortems
- Define and track SLOs/SLIs and error budgets in partnership with engineering and product teams
- Design and maintain observability and APM tooling (metrics, logs, tracing, dashboards, alerting) so issues are caught before they impact customers
- Continuously refine alert thresholds and runbooks to reduce noise and mean time to resolution
- Reduce operational toil through automation — scripting, self-healing systems, and repeatable processes
- Build and maintain CI/CD pipelines that let the development team ship safely and frequently
- Provision and manage infrastructure as code (Bicep) and follow standard change control and version control practices
- Ensure application infrastructure meets security and compliance requirements (e.g., SOC 2, GovRAMP) in partnership with the compliance team
- Apply security best practices to infrastructure design and change management
- Partner with the agile development team to translate business requirements into reliable technical solutions
- Participate in technical design sessions and produce clear documentation (diagrams, runbooks, architecture notes)
- Stay current on new Azure capabilities, industry standards, and SRE best practices, and bring recommendations back to the team
- Other duties as assigned
Skills
- Must be eligible to work in the U.S
- • Solid understanding of networking fundamentals, HTTP/S, and observability principles
- • Ability to evaluate multiple technical approaches and recommend the most effective solution for the context
- • Strong independent problem-solving skills balanced with effective collaboration in a team environment
- • Familiarity with software development lifecycle and programming/coding standards
- • Clear, professional communication, especially under incident pressure
- 3+ years of experience in an IT Operations, DevOps, or SRE role
- Hands-on technical experience with Microsoft Azure in a production environment
- Experience with infrastructure as code — Terraform and/or Bicep
- Proficiency in at least one scripting/programming language — Python or TypeScript
- Experience working with REST and/or GraphQL APIs
- Experience defining and tracking KPIs/SLOs for a web-based application
- Comfortable participating in an on-call rotation
- • Experience with compliance audits (SOC 2 Type 2, GovRAMP)
- • Familiarity with security frameworks (NIST, ISO 27001)
- • AZ-104 certification, or equivalent Azure networking experience
- • Experience with Node.js
- • Experience with low-code platforms (Power Apps, Logic Apps)
- • Familiarity with Scrum/Agile methodology and supporting tools (Confluence, JIRA, Git, Jenkins, Bamboo, TFS)
- • Ability to translate business requirements directly into application/site behavior changes
Qualifications
Must Haves
- must be eligible to work in the U.S
- • Solid understanding of networking fundamentals, HTTP/S, and observability principles
- • Ability to evaluate multiple technical approaches and recommend the most effective solution for the context
- • Strong independent problem-solving skills balanced with effective collaboration in a team environment
- • Familiarity with software development lifecycle and programming/coding standards
- • Clear, professional communication, especially under incident pressure
- 3+ years of experience in an IT Operations, DevOps, or SRE role
- Hands-on technical experience with Microsoft Azure in a production environment
- Experience with infrastructure as code — Terraform and/or Bicep
- Proficiency in at least one scripting/programming language — Python or TypeScript
- Experience working with REST and/or GraphQL APIs
- Experience defining and tracking KPIs/SLOs for a web-based application
- Comfortable participating in an on-call rotation
Nice to Haves
- • Experience with compliance audits (SOC 2 Type 2, GovRAMP)
- • Familiarity with security frameworks (NIST, ISO 27001)
- • AZ-104 certification, or equivalent Azure networking experience
- • Experience with Node.js
- • Experience with low-code platforms (Power Apps, Logic Apps)
- • Familiarity with Scrum/Agile methodology and supporting tools (Confluence, JIRA, Git, Jenkins, Bamboo, TFS)
- • Ability to translate business requirements directly into application/site behavior changes
Benefits
- 100% remote work
- Eligible for the company's commission plan
- Profit sharing