Restaurant365 logo
Restaurant365
Verified live 9h ago

Site Reliability Engineer II

Brief overview

Irvine, CAIn-person
UndergradOr in progress
$99k–$138k/yrStated range
2+ yrsMinimum
4 H-1B approvalsDept. of Labor
TerraformAnsibleCloudFormationPythonBashPowerShellLinux engineeringWindows administrationNginxApache TomcatGitLabGitPrometheusGrafanaELKSite24x7Nagios

About the company

Restaurant365 logo
Restaurant365restaurant365.com

SaaS company delivering restaurant accounting and back-office software

Visa sponsorship history

2 years sponsoring, last filed FY2025

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
4H-1B approved
100%approval rate
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20231
20242
20251

Job description

Summary

Restaurant365 is a SaaS company disrupting the restaurant industry with its cloud-based platform that centralizes accounting and back-office operations for restaurants. The Site Reliability Engineer II will support, enhance, and maintain the company's cloud infrastructure and applications, collaborating with various teams to improve reliability, scalability, and security of the SaaS platform.

Responsibilities

  • Respond to production incidents, perform triage and troubleshooting, and contribute to post-incident analysis
  • Identify and automate manual processes to improve efficiency and reduce risk
  • Enhance and evolve monitoring tools and platforms to improve observability
  • Promote and apply best practices for reliability, scalability, and performance across engineering
  • Implement and support cloud automation using Terraform, Ansible, or CloudFormation
  • Work within change management protocols to provide maximum uptime for production systems
  • Participate in on-call rotation, providing 24x7 support for incidents and contributing to root cause analysis
  • Partner with developers, architects, vendors, and IT teams to ensure reliable system operations
  • Research and remediate vulnerabilities in coordination with security teams
  • Maintain documentation of infrastructure, monitoring, runbooks, and incident response procedures
  • Apply company policies and procedures when handling operational tasks and incidents
  • Suggest and implement improvements to operational processes and monitoring practices
  • Contribute to technical diagrams, documentation, and runbooks for system reliability
  • Expand expertise in cloud services (Azure, AWS, or GCP) and container platforms (EKS, ECS, AKS)
  • Build proficiency with observability and monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios)
  • Develop scripting and automation skills using Python, Bash, PowerShell, or similar
  • Participate in planning discussions by contributing technical input on system stability and reliability

Skills

  • BS in Computer Science, Information Systems, or related field (or equivalent experience)
  • 2-4 years of experience in site reliability engineering, DevOps, or cloud operations
  • Experience with cloud platforms (Azure or AWS), including services such as AKS, ECS/EKS, Functions/Lambda, S3, and Blob storage
  • Proficiency with infrastructure-as-code and automation (Terraform, Ansible, YAML, Python, Bash, PowerShell)
  • Strong Linux engineering skills; working knowledge of Windows administration
  • Experience supporting production environments and participating in on-call rotations
  • Familiarity with web servers and middleware (Nginx, Apache Tomcat)
  • Experience with CI/CD tools (GitLab, Git, or similar)
  • Strong written, oral, and interpersonal communication skills
  • Experience with monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios)
  • Knowledge of performance analysis and system vulnerability remediation
  • Cloud certification (AWS or Azure) preferred
  • Familiarity with restaurant industry SaaS platforms and customer-facing applications

Qualifications

Must Haves

  • BS in Computer Science, Information Systems, or related field (or equivalent experience)
  • 2-4 years of experience in site reliability engineering, DevOps, or cloud operations
  • Experience with cloud platforms (Azure or AWS), including services such as AKS, ECS/EKS, Functions/Lambda, S3, and Blob storage
  • Proficiency with infrastructure-as-code and automation (Terraform, Ansible, YAML, Python, Bash, PowerShell)
  • Strong Linux engineering skills; working knowledge of Windows administration
  • Experience supporting production environments and participating in on-call rotations
  • Familiarity with web servers and middleware (Nginx, Apache Tomcat)
  • Experience with CI/CD tools (GitLab, Git, or similar)
  • Strong written, oral, and interpersonal communication skills

Nice to Haves

  • Experience with monitoring tools (Prometheus, Grafana, ELK, Site24x7, Nagios)
  • Knowledge of performance analysis and system vulnerability remediation
  • Cloud certification (AWS or Azure) preferred
  • Familiarity with restaurant industry SaaS platforms and customer-facing applications

Benefits

  • Comprehensive medical benefits, 100% paid for employee
  • 401k + matching
  • Equity Option Grant
  • Unlimited PTO + Company holidays
  • Wellness initiatives

More jobs like this