Summary
IMG Academy is a sports education brand providing on-campus and online experiences for student-athletes, coaches, and partners. The DevOps / Site Reliability Engineer will improve system reliability, scalability, developer productivity, and application performance by managing CI/CD pipelines, cloud infrastructure, monitoring, deployments, and incident response.
Responsibilities
- Configure, manage, and improve Bitbucket pipelines for deploying our applications to staging and production
- Improve CI pipeline speed, reliability, and security in collaboration with our Cloud Security Engineer
- Assist developers and QA teams with deployments
- Work with Docker and AWS ECR for container builds and deployment workflows
- Review and investigate system issues flagged by Sentry, NewRelic, and CloudWatch
- Monitor application performance, identify bottlenecks, and propose solutions
- Respond to production and staging issues, including database latency, unresponsive resources, or failed jobs
- Maintain and support non-production environments used by developers and QA
- Maintain and improve AWS infrastructure and Terraform resources
- Perform updates and upgrades to AWS services as needed to ensure reliability and ability to scale
- Partner with engineers to design systems that are scalable, observable, and resilient
- Work closely with our cloud security engineer to ensure secure configurations in CI/CD, AWS, and containerized workloads
- Contribute ideas and improvements to workflows, automation, and monitoring strategies
- Leverage AI to automate monitoring and diagnosis
Skills
- 3+ years of experience in DevOps, SRE, or related engineering roles
- Strong experience configuring CI/CD pipelines (Bitbucket Pipelines, GitHub Actions, or similar)
- Experience configuring, debugging and deploying PHP applications
- Hands-on experience with Docker and AWS ECR for container builds and deployments
- Strong experience with AWS services (EC2, RDS, ECS, Lambda, etc.) and Terraform for infrastructure as code
- Familiarity with monitoring and observability tools such as New Relic, Sentry, CloudWatch, or similar
- Strong troubleshooting skills for debugging performance issues in databases, applications, and distributed systems
- Experience with modern software development workflows (agile teams, code reviews, branching strategies)
- Strong scripting and automation skills (Bash, Python, or similar)
- Excellent communication skills and a collaborative mindset
- Interest in leveraging AI agents to automate monitoring and diagnosis workflows
Qualifications
Must Haves
- 3+ years of experience in DevOps, SRE, or related engineering roles
- Strong experience configuring CI/CD pipelines (Bitbucket Pipelines, GitHub Actions, or similar)
- Experience configuring, debugging and deploying PHP applications
- Hands-on experience with Docker and AWS ECR for container builds and deployments
- Strong experience with AWS services (EC2, RDS, ECS, Lambda, etc.) and Terraform for infrastructure as code
- Familiarity with monitoring and observability tools such as New Relic, Sentry, CloudWatch, or similar
- Strong troubleshooting skills for debugging performance issues in databases, applications, and distributed systems
- Experience with modern software development workflows (agile teams, code reviews, branching strategies)
- Strong scripting and automation skills (Bash, Python, or similar)
- Excellent communication skills and a collaborative mindset
- Interest in leveraging AI agents to automate monitoring and diagnosis workflows
Benefits
- Comprehensive Medical, Dental and Vision
- Flexible Spending Account and Health Savings Account options
- 401k with an Employer Match
- Short Term and Long Term Disability
- Group and Supplemental Life & AD&D
- Gym Discount Program
- Pet Insurance
- Wellbeing Program