CrewAI logo
CrewAI
Posted 24 days agoVerified live 1d ago

Software Engineer, Infrastructure & Reliability

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
AWSDockerCI/CDGitHub ActionsAmazon ECSKubernetesPostgreSQLRedisPythonIAMSecrets ManagementInfrastructure and Platform Engineering

About the company

CrewAI is a multi-agent automation platform for enterprises to automate complex tasks.

Job description

Summary

CrewAI is a framework and enterprise platform for building and orchestrating multi-agent AI systems. The Software Engineer, Infrastructure & Reliability will build and operate infrastructure across AWS, Azure, and GCP, focusing on containers, CI/CD, deployment automation, observability, security, and runtime reliability. The role also develops internal tooling and automation for scalable cloud and self-hosted deployments while reducing operational toil.

Responsibilities

  • Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services
  • Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety
  • Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs
  • Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior
  • Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems
  • Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access
  • Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time
  • Reduce operational toil by automating recurring workflows and making deployments boring

Skills

  • Strong infrastructure/platform engineering experience in production SaaS environments
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management
  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems
  • Experience supporting enterprise/self-hosted deployments
  • Terraform or other IaC experience
  • SRE background: SLOs, incident review, capacity planning, load testing
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability

Qualifications

Must Haves

  • Strong infrastructure/platform engineering experience in production SaaS environments
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management

Nice to Haves

  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems
  • Experience supporting enterprise/self-hosted deployments
  • Terraform or other IaC experience
  • SRE background: SLOs, incident review, capacity planning, load testing
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability

Benefits

  • Remote work arrangement

More jobs like this