Summary
Weekday AI is seeking an experienced Cloud / DevOps Engineer to support a cutting-edge GenAI environment. The role focuses on evaluating technical workflows, developing reference solutions, and establishing rigorous standards for AI-generated reasoning across cloud infrastructure, Kubernetes, AWS, Infrastructure-as-Code, and CI/CD.
Responsibilities
- Collaborate with research and engineering teams to identify knowledge gaps and improve AI model performance across cloud infrastructure, DevOps, Kubernetes, and Infrastructure-as-Code domains
- Design realistic and technically challenging tasks covering Kubernetes troubleshooting, AWS service integration, infrastructure automation, and production operations
- Develop accurate, detailed reference solutions for complex infrastructure engineering scenarios
- Review and evaluate AI-generated technical solutions for correctness, reliability, scalability, security, and adherence to production best practices
- Provide clear, structured written feedback highlighting technical gaps, incorrect assumptions, and opportunities for improvement
- Create detailed evaluation criteria, rubrics, and benchmarks for assessing Kubernetes troubleshooting, IaC architecture, AWS integrations, and CI/CD reasoning
- Develop scenarios involving cluster failures, infrastructure automation, deployment workflows, service integrations, and operational reliability
- Work closely with other technical subject matter experts to maintain consistency, accuracy, and quality across evaluation datasets
- Translate practical production experience into structured guidance that can be used to improve AI-generated infrastructure solutions
Skills
- **4+ years of professional experience** in Cloud Infrastructure, DevOps, Site Reliability Engineering, Platform Engineering, or a closely related field
- Strong hands-on experience managing **Kubernetes in production environments**, including diagnosing, troubleshooting, and resolving cluster failures and operational issues
- Experience with Kubernetes beyond simply writing manifests or consuming managed Kubernetes control planes
- Proven production experience with **Infrastructure-as-Code**, particularly **Terraform and/or AWS CDK**
- Strong practical knowledge of **AWS cloud services**, including production integration with services such as:
- AWS Lambda
- API Gateway
- DynamoDB
- Experience designing, implementing, and maintaining **CI/CD pipelines** for production workloads
- Strong understanding of cloud architecture, infrastructure automation, deployment strategies, observability, reliability, and operational best practices
- Demonstrated career progression with increasing ownership and responsibility in infrastructure, DevOps, or platform engineering
- Ability to commit reliably to **40 hours per week during standard weekdays**
- Excellent written and verbal communication skills, with the ability to explain complex technical concepts and engineering decisions clearly
- Strong analytical and troubleshooting abilities, particularly when diagnosing distributed systems and infrastructure failures
- Experience working with large-scale cloud infrastructure or highly distributed systems
- Familiarity with Kubernetes networking, security, storage, scaling, and cluster lifecycle management
- Experience implementing infrastructure security and reliability best practices
- Knowledge of AWS architecture patterns and cloud-native application design
- Experience with GitOps, containerization, monitoring, logging, and observability platforms
- Familiarity with modern DevOps and platform engineering methodologies
- Experience reviewing or evaluating technical documentation, engineering solutions, or AI-generated outputs
Qualifications
Must Haves
- **4+ years of professional experience** in Cloud Infrastructure, DevOps, Site Reliability Engineering, Platform Engineering, or a closely related field
- Strong hands-on experience managing **Kubernetes in production environments**, including diagnosing, troubleshooting, and resolving cluster failures and operational issues
- Experience with Kubernetes beyond simply writing manifests or consuming managed Kubernetes control planes
- Proven production experience with **Infrastructure-as-Code**, particularly **Terraform and/or AWS CDK**
- Strong practical knowledge of **AWS cloud services**, including production integration with services such as:
- AWS Lambda
- API Gateway
- DynamoDB
- Experience designing, implementing, and maintaining **CI/CD pipelines** for production workloads
- Strong understanding of cloud architecture, infrastructure automation, deployment strategies, observability, reliability, and operational best practices
- Demonstrated career progression with increasing ownership and responsibility in infrastructure, DevOps, or platform engineering
- Ability to commit reliably to **40 hours per week during standard weekdays**
- Excellent written and verbal communication skills, with the ability to explain complex technical concepts and engineering decisions clearly
- Strong analytical and troubleshooting abilities, particularly when diagnosing distributed systems and infrastructure failures
Nice to Haves
- Experience working with large-scale cloud infrastructure or highly distributed systems
- Familiarity with Kubernetes networking, security, storage, scaling, and cluster lifecycle management
- Experience implementing infrastructure security and reliability best practices
- Knowledge of AWS architecture patterns and cloud-native application design
- Experience with GitOps, containerization, monitoring, logging, and observability platforms
- Familiarity with modern DevOps and platform engineering methodologies
- Experience reviewing or evaluating technical documentation, engineering solutions, or AI-generated outputs