Summary
Tekmetric is a cloud-based software company providing solutions that help auto repair shops manage operations, payments, and customer engagement. The Site Reliability Engineer will design and maintain scalable, secure cloud infrastructure, improve system reliability and performance, automate deployment and infrastructure processes, and support disaster recovery, security, and compliance. The role also involves cross-functional collaboration and mentoring junior DevOps team members.
Responsibilities
- Design and implement scalable infrastructure: Architect and maintain reliable, scalable, and secure cloud infrastructure that supports positive user experiences and measurable business growth
- Monitor and optimize system performance: Develop and maintain monitoring, alerting, and incident response practices to ensure system reliability and performance at scale
- Automate everything: Create automated pipelines for deployment, testing, and infrastructure management to improve speed, consistency, and reliability across the organization
- Ensure high availability and disaster recovery: Implement and manage solutions for backup, disaster recovery, and failover processes to ensure business continuity
- Security and compliance: Apply best practices in security, monitoring, and compliance, ensuring that systems meet necessary requirements and regulations
- Collaboration: Work cross-functionally with development, data, product, and QA teams to improve application reliability and scalability
- Leadership and mentorship: Provide technical leadership, mentorship, and guidance to junior DevOps team members, fostering a culture of continuous learning and improvement
Skills
- 3+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related field, with deep knowledge of cloud environments (preferably AWS or GCP.)
- Hands-on experience with AWS (or similar cloud providers) and infrastructure as code (Terraform, etc.)
- Strong experience in automation tools
- Expertise in working with containerized environments like Docker and orchestration tools such as Kubernetes
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack)
- Proficiency in scripting languages like Python, Bash, or similar
- Experience with designing and optimizing Continuous Integration and Continuous Deployment (CI/CD) pipelines
- Strong communication skills and ability to work cross-functionally, solving complex technical challenges in a collaborative manner
- Ability to troubleshoot and resolve critical issues in high-pressure environments, maintaining composure and professionalism
- Experience with Infrastructure as Code tools like Terraform
- Familiarity with monitoring tools like Prometheus, Grafana, or the ELK stack
- Exposure to compliance and security best practices in cloud environments
- Experience coding in one or multiple programming languages such as Go, Java, Javascript
Qualifications
Must Haves
- 3+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related field, with deep knowledge of cloud environments (preferably AWS or GCP.)
- Hands-on experience with AWS (or similar cloud providers) and infrastructure as code (Terraform, etc.)
- Strong experience in automation tools
- Expertise in working with containerized environments like Docker and orchestration tools such as Kubernetes
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack)
- Proficiency in scripting languages like Python, Bash, or similar
- Experience with designing and optimizing Continuous Integration and Continuous Deployment (CI/CD) pipelines
- Strong communication skills and ability to work cross-functionally, solving complex technical challenges in a collaborative manner
- Ability to troubleshoot and resolve critical issues in high-pressure environments, maintaining composure and professionalism
Nice to Haves
- Experience with Infrastructure as Code tools like Terraform
- Familiarity with monitoring tools like Prometheus, Grafana, or the ELK stack
- Exposure to compliance and security best practices in cloud environments
- Experience coding in one or multiple programming languages such as Go, Java, Javascript
Benefits
- Hybrid and remote work models based on proximity to office hubs
- Travel to team and company-wide offsites several times a year is fully supported
- Competitive base salaries that reflect your value
- Generous Paid Time Off
- Paid maternity, parental bonding, and medical leave for you or your loved ones
- Comprehensive health benefits, including Medical, Dental, Vision, and Prescription coverage
- For employee-only plans, 100% of premiums are covered; 50% of family plan costs are covered
- Free, confidential counseling through BetterHelp
- 401(k) Retirement Savings Plan with 100% employer match on contributions up to 6%
- Flexible Spending Accounts (FSA) and Health Savings Accounts (HSA)
- Life and Accidental Death & Dismemberment (AD&D) Insurance
- Up to $60/month toward fitness, mental health, or almost anything that helps you feel your best
- After one year of employment, a $300 home office setup bonus
- Support for continuing education
- A stellar team of coworkers, a really cool office, and lots of fun activities