Zscaler logo
Zscaler
Posted 75 days agoVerified live 2d ago

Production Engineer

Brief overview

Remote
$102k–$128k/yrStated range
1+ yrsMinimum
428 H-1B approvalsDept. of Labor
69 green cardsCertified filings
PythonGoC++AWSGCPAzureLinuxRHELNetworking ProtocolsDistributed ArchitectureIncident ManagementITIL FrameworksPrometheusGrafanaOpenTelemetrySLI/SLO DefinitionError Budgets

About the company

Zscaler is a global cloud-based information security company that enables secure digital transformation for mobile and cloud.

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
428H-1B approved
99%approval rate
49new H-1B hires
69PERM certified
$187,741median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
202381
2024172
2025148
202627
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
202313
202444
202524
202620
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
202321
202410
202531
20267
Top sponsored roles
Principal Software Development EngineerStaff Software Development EngineerSr. Staff Software Development EngineerStaff Software EngineerSr. Staff Software Engineer
Sponsored employees from
IndiaChina

Job description

Summary

Zscaler is an AI-forward enterprise focused on digital transformation and cybersecurity. They are seeking a Production Engineer to enhance the reliability of their global platform while driving an automation-first culture across the company.

Responsibilities

  • Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments
  • Drive an "automation-first" culture by writing code (Python/Go) to eliminate manual toil and build self-healing systems
  • Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets
  • Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses
  • Partner with Engineering and partner teams to conduct operability reviews

Skills

  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • 1-3 years of experience managing reliability, scalability, and availability for large-scale production services
  • Deep expertise in programming (e.g., Python, Go, or C/C++)
  • Strong background in networking protocols, Linux/RHEL systems, and distributed architecture
  • Experience in high-stakes incident management and participation in a 24/7 on-call rotation
  • Proficiency in leveraging ITIL frameworks and incident data to drive service maturity through systematic problem management and technical operability reviews
  • Extensive experience with public cloud (AWS, Azure, GCP) and Infrastructure-as-Code (Ansible, Terraform, Helm, Temporal)
  • Experience with chaos engineering and disaster recovery planning at scale
  • Expertise in global routing (BGP) and traffic tunneling (GRE, IPSec) with a deep understanding of L7 proxy architectures (HAProxy), DNS at scale, and OS networking stack internals

Qualifications

Must Haves

  • Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • 1-3 years of experience managing reliability, scalability, and availability for large-scale production services
  • Deep expertise in programming (e.g., Python, Go, or C/C++)
  • Strong background in networking protocols, Linux/RHEL systems, and distributed architecture
  • Experience in high-stakes incident management and participation in a 24/7 on-call rotation
  • Proficiency in leveraging ITIL frameworks and incident data to drive service maturity through systematic problem management and technical operability reviews

Nice to Haves

  • Extensive experience with public cloud (AWS, Azure, GCP) and Infrastructure-as-Code (Ansible, Terraform, Helm, Temporal)
  • Experience with chaos engineering and disaster recovery planning at scale
  • Expertise in global routing (BGP) and traffic tunneling (GRE, IPSec) with a deep understanding of L7 proxy architectures (HAProxy), DNS at scale, and OS networking stack internals

Benefits

  • Various health plans
  • Time off plans for vacation and sick time
  • Parental leave options
  • Retirement options
  • Education reimbursement
  • In-office perks, and more!

More jobs like this