Russell Tobin logo
Russell Tobin
Posted 10 days agoVerified live 2d ago

Site Reliability Engineer - Remote

Brief overview

Remote
UndergradOr in progress
$55–$60/hrStated range
3+ yrsMinimum
Site Reliability EngineeringAWSKubernetesHelmLinuxAWS VPC NetworkingTerraformGitOpsDistributed Version ControlPrometheusGrafanaSoftware Configuration Management

About the company

Russell Tobin logo
Russell Tobinrusselltobin.com

Russell Tobin is a staffing and recruiting company that provides recruitment and staffing advisory services.

Job description

Summary

The confidential global client is a technology company seeking a Site Reliability Engineer. This role deploys and supports cloud-premises and SaaS software, responds to incidents, maintains infrastructure reliability, automates operational tasks, and provides advanced technical support and guidance to customers.

Responsibilities

  • Deploy software for Cloud Prem and SAAS customers
  • Respond to and diagnose system incidents in a timely and efficient manner, minimizing downtime and impact on users
  • Collaborate with other engineers to establish root causes and implement effective resolutions
  • Continuously improve incident response processes and documentation for future occurrences
  • Proactively monitor and maintain the health and performance of our infrastructure and services
  • Perform routine administrative tasks such as system configuration, user management, and data backups
  • Identify and implement operational improvements to ensure ongoing system reliability and efficiency
  • Develop and implement scripts and automated solutions to streamline operational tasks and reduce manual workload
  • Participate in the on-call rotation to address critical incidents outside of regular business hours
  • Ensure effective handoff between on-call engineers and document post-incident information for future reference
  • Document processes for support and create, maintain and execute run-books for identified situations
  • Provide tier 2/3 technical support to customers experiencing platform issues or requiring advanced troubleshooting
  • Work directly with customer technical teams to resolve complex deployment, configuration, and integration challenges
  • Conduct technical onboarding sessions and provide guidance on best practices for customer implementations
  • Collaborate with customer success teams to ensure smooth customer experiences and rapid issue resolution
  • Create and maintain customer-facing technical documentation, troubleshooting guides, and knowledge base articles
  • Escalate customer feedback and feature requests to product and engineering teams
  • Participate in customer calls and technical discussions to provide expert-level platform guidance
  • Track and analyze customer support metrics to identify trends and areas for improvement

Skills

  • 3+ years of experience in Site Reliability Engineering
  • 2+ years of experience working with cloud platform and cloud automation tools, especially in AWS
  • Strong experience with Kubernetes, Helm, Linux, AWS networking(VPC) and Terraform
  • Experience with the GitOps model for deployment
  • Familiarity with distributed version control
  • Experience with monitoring and alerting tools (e.g., Prometheus, Grafana)
  • Understanding of software configuration best practices
  • Ability to wear multiple hats in a fast-paced environment
  • Hands-on, “can do” attitude and a bias for action
  • Low ego and high intellectual curiosity
  • Comfortable working across time zones to support global customer base
  • Excellent communication skills with ability to explain technical concepts to both technical and non-technical audiences
  • Strong customer service orientation with patience and empathy when working with frustrated customers
  • Bazel and Cue Lang experience a plus

Qualifications

Must Haves

  • 3+ years of experience in Site Reliability Engineering
  • 2+ years of experience working with cloud platform and cloud automation tools, especially in AWS
  • Strong experience with Kubernetes, Helm, Linux, AWS networking(VPC) and Terraform
  • Experience with the GitOps model for deployment
  • Familiarity with distributed version control
  • Experience with monitoring and alerting tools (e.g., Prometheus, Grafana)
  • Understanding of software configuration best practices
  • Ability to wear multiple hats in a fast-paced environment
  • Hands-on, “can do” attitude and a bias for action
  • Low ego and high intellectual curiosity
  • Comfortable working across time zones to support global customer base
  • Excellent communication skills with ability to explain technical concepts to both technical and non-technical audiences
  • Strong customer service orientation with patience and empathy when working with frustrated customers

Nice to Haves

  • Bazel and Cue Lang experience a plus

Benefits

  • Eligible employees receive comprehensive healthcare coverage, including medical, dental, and vision plans.
  • Supplemental coverage, including accident insurance, critical illness insurance, and hospital indemnity.
  • 401(k) retirement savings.
  • Life and disability insurance.
  • Employee assistance program.
  • Legal support.
  • Auto insurance.
  • Home insurance.
  • Pet insurance.
  • Employee discounts with preferred vendors.
  • Remote work arrangement

More jobs like this