Storyteller logo
Storyteller
Posted 75 days agoVerified live 2d ago

Site Reliability / Production Engineer

Brief overview

Remote
Incident responseLog analysisMetrics analysisTracingDashboard monitoringAPI diagnosticsCode debuggingRunbook creationCloud platforms - AzureDistributed application diagnosticsDatabase managementQueue managementBackground processing systemsObservability platformsAlerting platformsIncident management platformsInfrastructure automation

About the company

We built a Stories SDK so you don't have to

Job description

Summary

Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. They are seeking two Site Reliability / Production Engineers to respond to live incidents, assess customer impact, and improve system reliability. The role involves working closely with support teams and product developers to ensure operational effectiveness and system improvements.

Responsibilities

  • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state
  • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code
  • Use AI throughout triage and diagnosis while checking its conclusions against real evidence
  • Choose and execute a proportionate mitigation, rollback, repair or bounded fix
  • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green
  • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication
  • Join customer conversations occasionally when direct technical involvement is genuinely useful
  • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix
  • Escalate with evidence, customer impact, actions already taken and the specific decision or help required
  • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait
  • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners
  • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes
  • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements
  • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve
  • Create safe, supervised automation for common operational actions
  • Work with product teams to close observability, rollback, runbook and supportability gaps
  • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner
  • Make reliability and on-call performance easier for the company to understand and improve over time

Skills

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it
  • Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response
  • Cloud platforms such as Azure or Cloudflare
  • Distributed application and API diagnostics
  • Databases, queues and background-processing systems
  • Observability, alerting and incident-management platforms
  • Infrastructure, deployment and release automation
  • Application development and safe production debugging
  • AI coding agents and workflow automation

Qualifications

Must Haves

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it

Nice to Haves

  • Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response
  • Cloud platforms such as Azure or Cloudflare
  • Distributed application and API diagnostics
  • Databases, queues and background-processing systems
  • Observability, alerting and incident-management platforms
  • Infrastructure, deployment and release automation
  • Application development and safe production debugging
  • AI coding agents and workflow automation

Benefits

  • 🌎 Fully remote working from anywhere in Tunisia! 
  • 🌙 Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty
  • The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed clearly during the hiring process.
  • We compensate you for completing it regardless of the outcome.

More jobs like this