Summary
VERSANT Media is an innovative media and entertainment company that operates across various core markets. They are seeking a System Reliability Operations (SRO) Engineer to establish and scale reliability practices within their technology organization, developing standards and documentation to enhance the reliability and supportability of their platforms.
Responsibilities
- Help establish, document, and maintain reliability engineering standards for applications, platforms, infrastructure, and shared services
- Translate SRO and SRE principles into practical guidelines, checklists, templates, and operating procedures that can be adopted across engineering teams
- Support the definition and implementation of service-level indicators, service-level objectives, availability targets, and operational health measures
- Develop standard criteria for production readiness, operational acceptance, service ownership, monitoring, escalation, and support
- Identify opportunities to standardize reliability practices across teams while accounting for different platforms, technologies, and business requirements
- Research operational trends and lessons learned from incidents and recommend updates to standards and documentation
- Create and maintain operational runbooks, troubleshooting guides, service support documentation, escalation procedures, and standard operating procedures
- Develop reusable templates for production readiness reviews, incident response, root cause analysis, service inventories, dependency mapping, and operational handoffs
- Partner with engineering teams to improve the quality, consistency, and accessibility of technical and operational documentation
- Help establish documentation ownership, review cycles, version control, and governance practices
- Organize operational knowledge in platforms such as Confluence, ServiceNow, or similar enterprise knowledge-management tools
- Ensure critical services have current documentation covering architecture, dependencies, monitoring, recovery procedures, support contacts, and known risks
- Promote documentation as an ongoing engineering responsibility rather than a one-time delivery activity
- Support operational readiness reviews for new services, major releases, migrations, and significant platform changes
- Validate that services have appropriate monitoring, alerting, dashboards, support procedures, dependency documentation, and recovery plans before production launch
- Work with engineering teams to identify and close readiness gaps
- Help define consistent service onboarding and operational acceptance processes for the Platform SRO organization
- Maintain service inventories and ownership information for critical enterprise platforms
- Support disaster recovery, failover, capacity, and resilience testing activities
- Participate in incident response for production and platform issues, supporting technical investigation, coordination, communications, and documentation
- Assist incident commanders and technical teams during high-severity events by maintaining timelines, tracking actions, and organizing relevant service information
- Support post-incident reviews and root cause analysis, ensuring findings, contributing factors, and corrective actions are clearly documented
- Track corrective and preventive actions through completion and escalate overdue or recurring reliability risks
- Analyze incident patterns and recurring issues to identify opportunities for improved documentation, automation, monitoring, or engineering standards
- Help maintain incident-response playbooks, severity definitions, escalation paths, and communication templates
- Partner with engineering teams to ensure applications and platforms have appropriate monitoring, logging, alerting, and dashboard coverage
- Help document observability requirements and recommended practices for enterprise services
- Review alerts for clarity, actionability, ownership, and alignment with documented response procedures
- Assist with the development of service-health dashboards and operational reporting
- Identify monitoring gaps and help teams create supporting runbooks and troubleshooting instructions
- Contribute to efforts that reduce alert noise and improve the quality of operational signals
- Identify repetitive operational tasks that can be standardized or automated
- Develop scripts, workflow automations, templates, or lightweight tools that improve documentation, readiness assessments, incident follow-up, and operational reporting
- Support integration of reliability checks and operational requirements into CI/CD and change-management workflows
- Help measure adoption of SRO standards and identify areas requiring additional guidance or enablement
- Contribute to the continuous improvement of Platform SRO processes, tools, and operating practices
- Partner with technical teams to understand their services, operational challenges, and support requirements
- Facilitate working sessions to create runbooks, document dependencies, define service objectives, and improve operational readiness
- Provide guidance to engineers on documentation, incident response, observability, and reliability fundamentals
- Develop internal reference materials and contribute to training or enablement sessions for engineering and operations teams
- Promote consistent reliability practices while building strong relationships across a matrixed enterprise organization
- Communicate operational risks, documentation gaps, and improvement recommendations clearly to technical stakeholders and Platform SRO leadership
Skills
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
- 3+ years of experience in System Reliability Operations, Site Reliability Engineering, IT Operations, Production Support, Platform Engineering, DevOps, Infrastructure Engineering, or a related technical role
- Experience supporting production applications, infrastructure, cloud platforms, or enterprise technology services
- Demonstrated experience creating technical documentation, operational runbooks, troubleshooting guides, or standard operating procedures
- Working knowledge of incident management, problem management, root cause analysis, and operational escalation practices
- Familiarity with reliability concepts such as SLIs, SLOs, availability, resiliency, service ownership, and operational readiness
- Experience with monitoring, logging, alerting, or observability tools
- Familiarity with cloud platforms such as AWS, Azure, or GCP and modern distributed systems
- Experience using enterprise collaboration and workflow tools such as ServiceNow, Jira, Confluence, or similar platforms
- Strong analytical and troubleshooting skills, with the ability to organize complex technical information into clear, actionable guidance
- Strong written and verbal communication skills with the ability to collaborate across engineering, infrastructure, security, and operations teams
- Ability to manage multiple documentation and operational improvement initiatives in a developing organization
- Experience supporting media, broadcast, streaming, digital publishing, or other highly available and time-sensitive environments
- Familiarity with ITIL Incident, Problem, Change, and Knowledge Management practices
- Experience supporting operational readiness reviews, disaster recovery exercises, or production service onboarding
- Familiarity with observability tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, Elastic, or similar platforms
- Experience with scripting or automation using Python, PowerShell, Bash, or comparable technologies
- Familiarity with CI/CD pipelines, Infrastructure as Code, Kubernetes, or containerized platforms
- Experience developing documentation standards, knowledge-management structures, or enterprise operating procedures
- Relevant cloud, ITIL, ServiceNow, or SRE-related certifications are a plus
Qualifications
Must Haves
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
- 3+ years of experience in System Reliability Operations, Site Reliability Engineering, IT Operations, Production Support, Platform Engineering, DevOps, Infrastructure Engineering, or a related technical role
- Experience supporting production applications, infrastructure, cloud platforms, or enterprise technology services
- Demonstrated experience creating technical documentation, operational runbooks, troubleshooting guides, or standard operating procedures
- Working knowledge of incident management, problem management, root cause analysis, and operational escalation practices
- Familiarity with reliability concepts such as SLIs, SLOs, availability, resiliency, service ownership, and operational readiness
- Experience with monitoring, logging, alerting, or observability tools
- Familiarity with cloud platforms such as AWS, Azure, or GCP and modern distributed systems
- Experience using enterprise collaboration and workflow tools such as ServiceNow, Jira, Confluence, or similar platforms
- Strong analytical and troubleshooting skills, with the ability to organize complex technical information into clear, actionable guidance
- Strong written and verbal communication skills with the ability to collaborate across engineering, infrastructure, security, and operations teams
- Ability to manage multiple documentation and operational improvement initiatives in a developing organization
Nice to Haves
- Experience supporting media, broadcast, streaming, digital publishing, or other highly available and time-sensitive environments
- Familiarity with ITIL Incident, Problem, Change, and Knowledge Management practices
- Experience supporting operational readiness reviews, disaster recovery exercises, or production service onboarding
- Familiarity with observability tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, Elastic, or similar platforms
- Experience with scripting or automation using Python, PowerShell, Bash, or comparable technologies
- Familiarity with CI/CD pipelines, Infrastructure as Code, Kubernetes, or containerized platforms
- Experience developing documentation standards, knowledge-management structures, or enterprise operating procedures
- Relevant cloud, ITIL, ServiceNow, or SRE-related certifications are a plus
Benefits
- Health insurance
- Retirement plans
- Paid time off