Summary
CreatorIQ is an operating system for creator-led growth that helps global brands and agencies transform creator marketing. The DevOps & AI/ML Infrastructure Engineer will support and improve cloud and ML/AI infrastructure, automate deployments, maintain CI/CD pipelines, and help ensure secure, scalable, and reliable development workflows. The role also supports MLOps platforms, observability, incident response, infrastructure security, and collaboration across engineering teams.
Responsibilities
- Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards
- Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation)
- Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations
- Support containerized environments and orchestration platforms
- Apply DevSecOps principles across infrastructure and deployment workflows
- Participate in disaster recovery planning, testing, and recovery activities
- Maintain and optimize CI/CD pipelines using tools such as GitLab CI/CD and Jenkins, supporting application and ML model deployments
- Improve deployment reliability and support zero-downtime deployment strategies
- Automate configuration management, infrastructure provisioning, and routine operational processes
- Troubleshoot deployment and pipeline issues and implement improvements to prevent recurrence
- Develop scripts and automation to reduce manual work and improve engineering efficiency
- Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases
- Use AI-assisted engineering tools, coding copilots, and AI-driven troubleshooting to improve DevOps productivity and reduce repetitive operational work
- Evaluate and adopt practical AI-enabled workflows that improve infrastructure management, troubleshooting, and operational efficiency
- Operate and scale ML platform infrastructure, including Databricks interactive clusters, jobs compute, ML pipelines, and Model Serving endpoints
- Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads
- Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health, while partnering with ML Engineering on model evaluation, quality thresholds, and model correctness
- Partner with ML Engineering to support reliable CI/CD and production deployment of ML models
- Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch
- Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues
- Improve system observability through effective log aggregation, metrics collection, monitoring, and alerting
- Partner with Software Engineers, ML Engineers, QA, and Software Engineers in Test to improve deployment workflows and integrate automated testing into CI/CD pipelines
- Collaborate with IT Security to maintain secure cloud operations and infrastructure policies
- Respond to engineering and Product Support requests in a timely manner and provide technical infrastructure support when needed
- Maintain accurate internal technical and operational documentation
- Collaborate effectively with international teams across multiple time zones
Skills
- 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role
- 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services
- 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS
- Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins
- Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies
- Strong Linux system administration and troubleshooting skills
- Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts
- Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks
- Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases
- Experience supporting data, ML, or other compute-intensive production workloads
- Experience with Google Cloud would be valuable, particularly for candidates who have worked across multi-cloud environments
- Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools would be beneficial
- Experience with serverless and event-driven architectures using technologies such as AWS Lambda, API Gateway, and SQS is a plus
- Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, firewalls, or similar technologies, would be valuable
- Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial
- Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions is a plus
- FinOps experience, including cloud cost monitoring, optimization, and accountability practices, would be valuable
- Experience with API gateways or API management platforms such as Kong, Apigee, or similar technologies is beneficial
- Experience with MLOps platforms and practices—particularly Databricks, model serving, ML pipelines, and model monitoring—would be an advantage
Qualifications
Must Haves
- 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role
- 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services
- 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS
- Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins
- Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies
- Strong Linux system administration and troubleshooting skills
- Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts
- Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks
- Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases
- Experience supporting data, ML, or other compute-intensive production workloads
Nice to Haves
- Experience with Google Cloud would be valuable, particularly for candidates who have worked across multi-cloud environments
- Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools would be beneficial
- Experience with serverless and event-driven architectures using technologies such as AWS Lambda, API Gateway, and SQS is a plus
- Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, firewalls, or similar technologies, would be valuable
- Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial
- Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions is a plus
- FinOps experience, including cloud cost monitoring, optimization, and accountability practices, would be valuable
- Experience with API gateways or API management platforms such as Kong, Apigee, or similar technologies is beneficial
- Experience with MLOps platforms and practices—particularly Databricks, model serving, ML pipelines, and model monitoring—would be an advantage
Benefits
- Flexible work model that combines both in-person and remote work
- Utilize our learning platform to fully get the training and tools you'll need to become successful here from your first day with us.
- 15 days of vacation, floating and company holidays, wellness benefits, and paid parental leave.
- Comprehensive medical, dental, vision, life, and disability insurance, plus additional wellness benefits.
- A 401(k) plan to help you plan ahead.
- Work from home stipend to assist you in setting up a home office that works for you.