Summary
AWeber is a remote-first company providing marketing communication software that helps small businesses build customer connections and grow. The Operations Engineer II will design and build solutions for AI agents, review agent-generated infrastructure code, and help keep the operational platform reliable, secure, and fast. The role includes infrastructure provisioning, monitoring and alerting, production support, release coordination, and participation in a 24x7 on-call rotation.
Responsibilities
- Designing and building solutions for AI agents to implement
- Reviewing agent-generated configuration and code for quality, correctness, and maintainability
- Coordinating that work across systems
- Provisioning infrastructure
- Tuning monitoring and alerting
- Participating in our 24x7 on-call rotation
- Defining infrastructure as code in Terraform
- Reviewing an agent's proposed Kubernetes change before it ships
- Wiring up a new alert in Grafana so the next incident gets caught sooner
- Generating documentation that’s usable by both humans and AI agents
- Working a release through GitHub and our deployment pipeline
- Manage context switches and competing priorities without losing quality or your footing
- Care about outcomes over output, and hold yourself to delivering results, not just tickets closed
- Have a growth mindset — you learn from experience and iterate on how you work, including how you work with AI
- Are self-motivated to adopt AI-assisted development practices rather than waiting to be told to
- Have a bias for action and real ownership mentality, balanced with the humility to ask for help
- Stay resilient through setbacks, shifting priorities, and ambiguity, without letting quality slip
- Connect your day-to-day work to customer impact, and bring that perspective to your teammates
- Are comfortable working autonomously on small-to-medium projects, while staying highly collaborative with the rest of the team
- Experience in a 24x7 on-call rotation (incident response, mitigation, root cause analysis), troubleshooting across Linux and Mac-based environments, and using Git or Mercurial for version control
- Experience with containerized, cloud-based infrastructure — Kubernetes, Docker, and infrastructure-as-code tools like Terraform or CloudFormation on AWS, GCP, or similar
- Comfort with shell scripting and Python, and experience configuring monitoring and alerting with tools like Grafana, Prometheus, PagerDuty, or Sentry
- Solid grasp of networking fundamentals (DNS, VPN, load balancing, subnetting) and security concepts (firewalls, secrets management, IAM, PKI), with practical experience applying them
Skills
- * Are detail‑oriented and genuinely curious — you want to know *why* something broke, not just that it's fixed
- * Can manage context switches and competing priorities without losing quality or your footing
- * Care about outcomes over output, and hold yourself to delivering results, not just tickets closed
- * Have a growth mindset — you learn from experience and iterate on how you work, including how you work with AI
- * Are self‑motivated to adopt AI‑assisted development practices rather than waiting to be told to
- * Have a bias for action and real ownership mentality, balanced with the humility to ask for help
- * Stay resilient through setbacks, shifting priorities, and ambiguity, without letting quality slip
- * Connect your day‑to‑day work to customer impact, and bring that perspective to your teammates
- * Are comfortable working autonomously on small‑to‑medium projects, while staying highly collaborative with the rest of the team
- * Bachelor's degree in a STEM‑related field or equivalent work experience, plus 2+ years in a platform engineering or systems/operations role
- * Working knowledge of AI development tools and coding assistants, with comfort using them in day‑to‑day implementation work
- * Strong analytical and troubleshooting skills — you fix root causes, not just symptoms — paired with the communication skills to translate technical issues for different audiences
- * Experience in a 24x7 on‑call rotation (incident response, mitigation, root cause analysis), troubleshooting across Linux and Mac‑based environments, and using Git or Mercurial for version control
- * Experience with containerized, cloud‑based infrastructure — Kubernetes, Docker, and infrastructure‑as‑code tools like Terraform or CloudFormation on AWS, GCP, or similar
- * Comfort with shell scripting and Python, and experience configuring monitoring and alerting with tools like Grafana, Prometheus, PagerDuty, or Sentry
- * Solid grasp of networking fundamentals (DNS, VPN, load balancing, subnetting) and security concepts (firewalls, secrets management, IAM, PKI), with practical experience applying them
Qualifications
Must Haves
- * Are detail‑oriented and genuinely curious — you want to know *why* something broke, not just that it's fixed
- * Can manage context switches and competing priorities without losing quality or your footing
- * Care about outcomes over output, and hold yourself to delivering results, not just tickets closed
- * Have a growth mindset — you learn from experience and iterate on how you work, including how you work with AI
- * Are self‑motivated to adopt AI‑assisted development practices rather than waiting to be told to
- * Have a bias for action and real ownership mentality, balanced with the humility to ask for help
- * Stay resilient through setbacks, shifting priorities, and ambiguity, without letting quality slip
- * Connect your day‑to‑day work to customer impact, and bring that perspective to your teammates
- * Are comfortable working autonomously on small‑to‑medium projects, while staying highly collaborative with the rest of the team
- * Bachelor's degree in a STEM‑related field or equivalent work experience, plus 2+ years in a platform engineering or systems/operations role
- * Working knowledge of AI development tools and coding assistants, with comfort using them in day‑to‑day implementation work
- * Strong analytical and troubleshooting skills — you fix root causes, not just symptoms — paired with the communication skills to translate technical issues for different audiences
- * Experience in a 24x7 on‑call rotation (incident response, mitigation, root cause analysis), troubleshooting across Linux and Mac‑based environments, and using Git or Mercurial for version control
- * Experience with containerized, cloud‑based infrastructure — Kubernetes, Docker, and infrastructure‑as‑code tools like Terraform or CloudFormation on AWS, GCP, or similar
- * Comfort with shell scripting and Python, and experience configuring monitoring and alerting with tools like Grafana, Prometheus, PagerDuty, or Sentry
- * Solid grasp of networking fundamentals (DNS, VPN, load balancing, subnetting) and security concepts (firewalls, secrets management, IAM, PKI), with practical experience applying them
Benefits
- 100% Remote
- Strong culture that supports flexibility, entrepreneurialism, and collaboration.
- 100% Company Paid PPO medical, dental, vision insurance (spousal and domestic partner benefits available).
- 4-7 weeks of paid time off and holidays (based on tenure).
- 4 week paid sabbatical (based on tenure).
- 401K retirement plan with 4% company match.
- Company Profit Share.
- Home office equipment and internet stipend.
- Tuition reimbursement, conferences, and learning opportunities.
- Gym Memberships Reimbursement.
- Company Paid Short Term Disability Insurance.
- Company Paid Life Insurance.
- Additional Supplemental Benefits (Long Term Disability, Critical Illness, and Additional Life Insurance).