Summary
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. As a Site Reliability Engineer, you will improve the reliability, scalability, and operational maturity of the Filevine platform, designing automation and solving production challenges to ensure seamless operations for legal teams.
Responsibilities
- Design, build, and maintain the monitoring, logging, distributed tracing, dashboards, and alerting that give teams meaningful visibility into production health
- Build automation, tooling, and CI/CD improvements that increase engineering efficiency, reduce toil, and support reliable deployments at scale
- Design, implement, and maintain reliable systems for building, deploying, testing, and operating Filevine products — proactively identifying and resolving reliability, performance, scalability, and security risks before they impact customers
- Participate in a shared 24/7 on-call rotation, using operational insights to drive automation and long-term reliability improvements; continuously improve runbooks, documentation, and engineering standards
- Take ownership of technical initiatives from design through implementation, develop deep expertise in critical areas of the Filevine platform, and communicate clearly with technical and business stakeholders
Skills
- 4+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles, including at least 2 years in a Site Reliability Engineering or reliability-focused role
- Working knowledge of distributed systems and how applications, infrastructure, and cloud services interact in production; demonstrated ability to troubleshoot production issues, perform root cause analysis, and drive long-term reliability improvements
- Proficiency with Python, Bash, or similar scripting languages; experience building production tooling, automation, or CI/CD pipelines and deployment automation
- Hands-on experience operating Kubernetes-based workloads and cloud infrastructure in AWS or a comparable platform, including compute, container orchestration, networking, IAM, object storage, and cloud-native monitoring
- Experience with Infrastructure as Code tools such as Terraform, Pulumi, or AWS CloudFormation, and familiarity with modern observability practices including monitoring, logging, alerting, distributed tracing, and incident response
- Experience using AI-assisted engineering tools to improve productivity, accelerate troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for building reliable systems through continuous improvement
- Strong written and verbal communication skills
- Bachelor's degree in Computer Science, Information Systems, or a related field, equivalent industry certifications, or comparable professional experience
Qualifications
Must Haves
- 4+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles, including at least 2 years in a Site Reliability Engineering or reliability-focused role
- Working knowledge of distributed systems and how applications, infrastructure, and cloud services interact in production; demonstrated ability to troubleshoot production issues, perform root cause analysis, and drive long-term reliability improvements
- Proficiency with Python, Bash, or similar scripting languages; experience building production tooling, automation, or CI/CD pipelines and deployment automation
- Hands-on experience operating Kubernetes-based workloads and cloud infrastructure in AWS or a comparable platform, including compute, container orchestration, networking, IAM, object storage, and cloud-native monitoring
- Experience with Infrastructure as Code tools such as Terraform, Pulumi, or AWS CloudFormation, and familiarity with modern observability practices including monitoring, logging, alerting, distributed tracing, and incident response
- Experience using AI-assisted engineering tools to improve productivity, accelerate troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for building reliable systems through continuous improvement
- Strong written and verbal communication skills
- Bachelor's degree in Computer Science, Information Systems, or a related field, equivalent industry certifications, or comparable professional experience
Benefits
- Medical, Dental, & Vision Insurance (for full-time employees)
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
- A dynamic, rapidly growing company, focused on helping organizations thrive