Storyteller logo
Storyteller
Posted 71 days agoVerified live 1d ago

Site Reliability / Production Engineer

Brief overview

Remote
Incident ResponseLog AnalysisMetrics AnalysisTracingDashboard MonitoringDeployment History AnalysisInfrastructure ManagementDatabase ManagementQueue ManagementAPI DiagnosticsApplication Code DebuggingRollback DeploymentsConfiguration ManagementData RepairCode Change DeploymentAlert TuningRunbook Creation

About the company

We built a Stories SDK so you don't have to

Job description

Summary

Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. They are seeking hands-on production engineers who can take ownership of live systems, investigate incidents, and coordinate responses to ensure reliability and improve systems after incidents.

Responsibilities

  • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state
  • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code
  • Use AI throughout triage and diagnosis while checking its conclusions against real evidence
  • Choose and execute a proportionate mitigation, rollback, repair or bounded fix
  • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green
  • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication
  • Join customer conversations occasionally when direct technical involvement is genuinely useful
  • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix
  • Escalate with evidence, customer impact, actions already taken and the specific decision or help required
  • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait
  • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners
  • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes
  • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements
  • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve
  • Create safe, supervised automation for common operational actions
  • Work with product teams to close observability, rollback, runbook and supportability gaps
  • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner
  • Make reliability and on-call performance easier for the company to understand and improve over time

Skills

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it
  • Cloud platforms such as Azure or Cloudflare
  • Distributed application and API diagnostics
  • Databases, queues and background-processing systems
  • Observability, alerting and incident-management platforms
  • Infrastructure, deployment and release automation
  • Application development and safe production debugging
  • AI coding agents and workflow automation

Qualifications

Must Haves

  • Agency and ownership - You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed
  • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded
  • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan
  • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions
  • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals
  • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work
  • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand
  • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it

Nice to Haves

  • Cloud platforms such as Azure or Cloudflare
  • Distributed application and API diagnostics
  • Databases, queues and background-processing systems
  • Observability, alerting and incident-management platforms
  • Infrastructure, deployment and release automation
  • Application development and safe production debugging
  • AI coding agents and workflow automation

Benefits

  • Fully remote working from anywhere in Egypt!
  • Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty
  • When there are no live incidents, the active shift will be used for reliability-improvement work.
  • The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed clearly during the hiring process.
  • Paid Take-home Task (~60-90 mins): We compensate you for completing it regardless of the outcome.

More jobs like this