Versant Media logo
Versant Media
Posted 60 days agoVerified live 1d ago

SRO Engineer

Brief overview

Remote
UndergradOr in progress
$130k–$160k/yrStated range
3+ yrsMinimum
System Reliability OperationsSite Reliability EngineeringIT OperationsProduction SupportPlatform EngineeringDevOpsInfrastructure EngineeringOperational RunbooksTroubleshooting GuidesStandard Operating ProceduresIncident ManagementProblem ManagementRoot Cause AnalysisOperational Escalation PracticesService-Level Indicators (SLIs)Service-Level Objectives (SLOs)Availability Targets

About the company

Versant Media logo
Versant Mediaversantmedia.com

Versant Media is a media company that offers television and digital entertainment services.

Visa sponsorship history

1 year sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
$204,118median wage / yr
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
202611
Top sponsored roles
Senior Site Reliability EngineerSenior Manager, AndroidDirector of Software EngineeringPrincipal DevSecOps EngineerSenior Manager, Core Services

Job description

Summary

VERSANT Media is an innovative media and entertainment company that operates across various core markets. They are seeking a System Reliability Operations (SRO) Engineer to establish and scale reliability practices within their technology organization, developing standards and documentation to enhance the reliability and supportability of their platforms.

Responsibilities

  • Help establish, document, and maintain reliability engineering standards for applications, platforms, infrastructure, and shared services
  • Translate SRO and SRE principles into practical guidelines, checklists, templates, and operating procedures that can be adopted across engineering teams
  • Support the definition and implementation of service-level indicators, service-level objectives, availability targets, and operational health measures
  • Develop standard criteria for production readiness, operational acceptance, service ownership, monitoring, escalation, and support
  • Identify opportunities to standardize reliability practices across teams while accounting for different platforms, technologies, and business requirements
  • Research operational trends and lessons learned from incidents and recommend updates to standards and documentation
  • Create and maintain operational runbooks, troubleshooting guides, service support documentation, escalation procedures, and standard operating procedures
  • Develop reusable templates for production readiness reviews, incident response, root cause analysis, service inventories, dependency mapping, and operational handoffs
  • Partner with engineering teams to improve the quality, consistency, and accessibility of technical and operational documentation
  • Help establish documentation ownership, review cycles, version control, and governance practices
  • Organize operational knowledge in platforms such as Confluence, ServiceNow, or similar enterprise knowledge-management tools
  • Ensure critical services have current documentation covering architecture, dependencies, monitoring, recovery procedures, support contacts, and known risks
  • Promote documentation as an ongoing engineering responsibility rather than a one-time delivery activity
  • Support operational readiness reviews for new services, major releases, migrations, and significant platform changes
  • Validate that services have appropriate monitoring, alerting, dashboards, support procedures, dependency documentation, and recovery plans before production launch
  • Work with engineering teams to identify and close readiness gaps
  • Help define consistent service onboarding and operational acceptance processes for the Platform SRO organization
  • Maintain service inventories and ownership information for critical enterprise platforms
  • Support disaster recovery, failover, capacity, and resilience testing activities
  • Participate in incident response for production and platform issues, supporting technical investigation, coordination, communications, and documentation
  • Assist incident commanders and technical teams during high-severity events by maintaining timelines, tracking actions, and organizing relevant service information
  • Support post-incident reviews and root cause analysis, ensuring findings, contributing factors, and corrective actions are clearly documented
  • Track corrective and preventive actions through completion and escalate overdue or recurring reliability risks
  • Analyze incident patterns and recurring issues to identify opportunities for improved documentation, automation, monitoring, or engineering standards
  • Help maintain incident-response playbooks, severity definitions, escalation paths, and communication templates
  • Partner with engineering teams to ensure applications and platforms have appropriate monitoring, logging, alerting, and dashboard coverage
  • Help document observability requirements and recommended practices for enterprise services
  • Review alerts for clarity, actionability, ownership, and alignment with documented response procedures
  • Assist with the development of service-health dashboards and operational reporting
  • Identify monitoring gaps and help teams create supporting runbooks and troubleshooting instructions
  • Contribute to efforts that reduce alert noise and improve the quality of operational signals
  • Identify repetitive operational tasks that can be standardized or automated
  • Develop scripts, workflow automations, templates, or lightweight tools that improve documentation, readiness assessments, incident follow-up, and operational reporting
  • Support integration of reliability checks and operational requirements into CI/CD and change-management workflows
  • Help measure adoption of SRO standards and identify areas requiring additional guidance or enablement
  • Contribute to the continuous improvement of Platform SRO processes, tools, and operating practices
  • Partner with technical teams to understand their services, operational challenges, and support requirements
  • Facilitate working sessions to create runbooks, document dependencies, define service objectives, and improve operational readiness
  • Provide guidance to engineers on documentation, incident response, observability, and reliability fundamentals
  • Develop internal reference materials and contribute to training or enablement sessions for engineering and operations teams
  • Promote consistent reliability practices while building strong relationships across a matrixed enterprise organization
  • Communicate operational risks, documentation gaps, and improvement recommendations clearly to technical stakeholders and Platform SRO leadership

Skills

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
  • 3+ years of experience in System Reliability Operations, Site Reliability Engineering, IT Operations, Production Support, Platform Engineering, DevOps, Infrastructure Engineering, or a related technical role
  • Experience supporting production applications, infrastructure, cloud platforms, or enterprise technology services
  • Demonstrated experience creating technical documentation, operational runbooks, troubleshooting guides, or standard operating procedures
  • Working knowledge of incident management, problem management, root cause analysis, and operational escalation practices
  • Familiarity with reliability concepts such as SLIs, SLOs, availability, resiliency, service ownership, and operational readiness
  • Experience with monitoring, logging, alerting, or observability tools
  • Familiarity with cloud platforms such as AWS, Azure, or GCP and modern distributed systems
  • Experience using enterprise collaboration and workflow tools such as ServiceNow, Jira, Confluence, or similar platforms
  • Strong analytical and troubleshooting skills, with the ability to organize complex technical information into clear, actionable guidance
  • Strong written and verbal communication skills with the ability to collaborate across engineering, infrastructure, security, and operations teams
  • Ability to manage multiple documentation and operational improvement initiatives in a developing organization
  • Experience supporting media, broadcast, streaming, digital publishing, or other highly available and time-sensitive environments
  • Familiarity with ITIL Incident, Problem, Change, and Knowledge Management practices
  • Experience supporting operational readiness reviews, disaster recovery exercises, or production service onboarding
  • Familiarity with observability tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, Elastic, or similar platforms
  • Experience with scripting or automation using Python, PowerShell, Bash, or comparable technologies
  • Familiarity with CI/CD pipelines, Infrastructure as Code, Kubernetes, or containerized platforms
  • Experience developing documentation standards, knowledge-management structures, or enterprise operating procedures
  • Relevant cloud, ITIL, ServiceNow, or SRE-related certifications are a plus

Qualifications

Must Haves

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
  • 3+ years of experience in System Reliability Operations, Site Reliability Engineering, IT Operations, Production Support, Platform Engineering, DevOps, Infrastructure Engineering, or a related technical role
  • Experience supporting production applications, infrastructure, cloud platforms, or enterprise technology services
  • Demonstrated experience creating technical documentation, operational runbooks, troubleshooting guides, or standard operating procedures
  • Working knowledge of incident management, problem management, root cause analysis, and operational escalation practices
  • Familiarity with reliability concepts such as SLIs, SLOs, availability, resiliency, service ownership, and operational readiness
  • Experience with monitoring, logging, alerting, or observability tools
  • Familiarity with cloud platforms such as AWS, Azure, or GCP and modern distributed systems
  • Experience using enterprise collaboration and workflow tools such as ServiceNow, Jira, Confluence, or similar platforms
  • Strong analytical and troubleshooting skills, with the ability to organize complex technical information into clear, actionable guidance
  • Strong written and verbal communication skills with the ability to collaborate across engineering, infrastructure, security, and operations teams
  • Ability to manage multiple documentation and operational improvement initiatives in a developing organization

Nice to Haves

  • Experience supporting media, broadcast, streaming, digital publishing, or other highly available and time-sensitive environments
  • Familiarity with ITIL Incident, Problem, Change, and Knowledge Management practices
  • Experience supporting operational readiness reviews, disaster recovery exercises, or production service onboarding
  • Familiarity with observability tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, Elastic, or similar platforms
  • Experience with scripting or automation using Python, PowerShell, Bash, or comparable technologies
  • Familiarity with CI/CD pipelines, Infrastructure as Code, Kubernetes, or containerized platforms
  • Experience developing documentation standards, knowledge-management structures, or enterprise operating procedures
  • Relevant cloud, ITIL, ServiceNow, or SRE-related certifications are a plus

Benefits

  • Health insurance
  • Retirement plans
  • Paid time off

More jobs like this