VITURE logo
VITURE
Posted 96 days agoVerified live 6h ago

Site Reliability Engineer

Brief overview

Cupertino, CAIn-person
UndergradOr in progress
$135k–$180k/yrStated range
2+ yrsMinimum

About the company

VITURE is a 2C technology company dedicated to developing AR technology, creating trendy XR glasses.

Job description

Summary

VITURE is the #1 XR glasses brand in the US, aiming to create the first great AI interface you wear. The Site Reliability Engineer will design and maintain scalable cloud infrastructure for intelligent eyewear, ensuring high reliability and performance for users worldwide.

Responsibilities

  • Design, deploy, and manage scalable infrastructure across mainstream cloud platforms to support our high-traffic AI services (e.g., LLM inference pipelines, real-time voice, and spatial computing backends)
  • Establish and execute incident management protocols, participate in on-call rotations, and lead blameless post-mortems to continuously reduce Mean Time to Recovery (MTTR) and improve system reliability
  • Champion Infrastructure as Code (IaC) principles. Leverage automation tools (e.g., Terraform, Ansible) to automate provisioning, configuration, and deployments, actively eliminating manual operational toil
  • Build and maintain comprehensive observability platforms (monitoring, logging, tracing) to track real-time resource utilization, define critical metrics, and ensure strict Service Level Objectives (SLOs) are met
  • Collaborate closely with software and AI teams to embed reliability early in the development lifecycle, streamline CI/CD pipelines, and optimize performance across servers and containerized environments
  • Enforce robust security postures by implementing network policies, configuring firewalls, and establishing rigorous backup and Disaster Recovery (DR) strategies
  • Mentor junior engineers and contribute to fostering a culture of engineering excellence and reliability as needed

Skills

  • Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field; or equivalent practical experience with a demonstrated track record of exceptional work
  • 2+ years of professional experience in an SRE, DevOps, or Cloud Infrastructure role, preferably supporting high-concurrency consumer applications or AI services
  • Strong hands-on experience managing and scaling distributed systems on mainstream public or private cloud ecosystems
  • Solid expertise in containerization (e.g., Docker) and orchestration platforms (e.g., Kubernetes) for deploying complex microservices architectures
  • Proficiency in scripting or programming languages (such as Python, Go, or Shell) to develop custom automation and tooling
  • Experience with modern observability, telemetry, and centralized logging stacks (e.g., Prometheus, Grafana, ELK Stack)
  • Excellent communication skills with the ability to articulate technical decisions and collaborate effectively with cross-functional teams
  • Previous experience managing infrastructure for AI/ML workloads (e.g., GPU cluster management, multimodal AI inference deployments)
  • Personal project experience is a strong plus — we love seeing what you build for fun. Share your side projects, open-source infrastructure contributions, or passion builds with us
  • Product-minded with a refined standard for reliability and craft — you care deeply about the user experience when systems degrade gracefully, not just whether the servers are running

Qualifications

Must Haves

  • Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field; or equivalent practical experience with a demonstrated track record of exceptional work
  • 2+ years of professional experience in an SRE, DevOps, or Cloud Infrastructure role, preferably supporting high-concurrency consumer applications or AI services
  • Strong hands-on experience managing and scaling distributed systems on mainstream public or private cloud ecosystems
  • Solid expertise in containerization (e.g., Docker) and orchestration platforms (e.g., Kubernetes) for deploying complex microservices architectures
  • Proficiency in scripting or programming languages (such as Python, Go, or Shell) to develop custom automation and tooling
  • Experience with modern observability, telemetry, and centralized logging stacks (e.g., Prometheus, Grafana, ELK Stack)
  • Excellent communication skills with the ability to articulate technical decisions and collaborate effectively with cross-functional teams

Nice to Haves

  • Previous experience managing infrastructure for AI/ML workloads (e.g., GPU cluster management, multimodal AI inference deployments)
  • Personal project experience is a strong plus — we love seeing what you build for fun. Share your side projects, open-source infrastructure contributions, or passion builds with us
  • Product-minded with a refined standard for reliability and craft — you care deeply about the user experience when systems degrade gracefully, not just whether the servers are running

Benefits

  • Discretionary bonus
  • Equity
  • Comprehensive health, dental, and vision insurance
  • 401(k) with company match
  • Flexible schedule
  • PTO, sick leave, and parental leave
  • Collaborative, inclusive, engineering-driven culture

More jobs like this