Summary
Fexa builds facilities-management software for companies with large retail footprints, focusing on reliability and customer impact. The Platform Engineer II will own components and projects end-to-end, design and deliver solutions, and contribute to the team's technical direction while ensuring operational excellence.
Responsibilities
- Author reusable Terraform modules that other engineers build on, and make changes confidently across our AWS infrastructure within our established conventions
- Design, optimize, and standardize continuous integration and deployment pipelines to accelerate time-to-market
- Own components and projects end-to-end, while taking on larger initiatives such as database engine migrations, cronjobs to EventBridge, or data warehouse infrastructure, shaping the approach with direction and support
- Contribute to Architecture Decisions and help refine our Terraform and observability standards
- Instrument systems, build dashboards and alerts, and begin defining the observability strategy for the systems you work on
- Apply security by default, including least-privilege IAM, proper secrets handling, and key and certificate rotation
- Carry a share of the team's day-to-day operational load alongside project work, including production support, escalations, maintenance, and keeping existing systems healthy. This role is not projects alone
- Write the runbooks that let others operate what you build, create documentation for both the team and the broader engineering organization, and help onboard and mentor newer engineers
- You will take part in the team's on-call rotation as an escalation point of contact. You will step in when first-line remediation and runbooks haven't resolved an incident, and you will drive the issue to resolution
Skills
- Bachelor's degree in Computer Science, Engineering, or a related field with 5+ years of relevant experience, or an equivalent combination of education and 8+ years of professional experience in platform engineering, infrastructure, or DevOps
- 5+ years in platform engineering, infrastructure, or DevOps roles, with a proven track record of building internal tooling and Production support
- Strong hands-on experience across AWS infrastructure
- Advanced Terraform: composing and versioning reusable modules; remote state with locking and workspaces/backends; meta-arguments and dynamic blocks (for_each, count, dynamic); safe state operations and refactoring (import, moved blocks, targeted state changes); drift detection; and policy/testing in CI (e.g., Terratest, tflint, Sentinel or OPA)
- Building and debugging CI/CD pipelines independently (GitHub Actions and/or Jenkins)
- Container orchestration on EKS/Kubernetes and AWS Fargate
- GitOps workflows (e.g., ArgoCD)
- Observability in practice: instrumenting services and building meaningful dashboards and alerts across metrics, logs, and APM
- Security fundamentals applied by default: least-privilege IAM, secrets handling, key and certificate rotation
- Strong proficiency with python and shell scripting to automate recurring work and build internal tooling
- Clear written and verbal communication — able to plan a piece of work, document the reasoning, and align with the team before building
- AWS certification, such as Solutions Architect Associate/Professional or SysOps/DevOps Engineer
- Ruby or Postgres performance tuning experience
- FinOps or cloud-cost optimization: right-sizing, cost tagging, and reporting
- OpenSearch/Elasticsearch operation and migration
- Data warehouse / analytics infrastructure on AWS (Redshift, Glue, Step Functions)
- Disaster recovery and backup/restore testing; exposure to RTO/RPO planning
- Experience migrating between CI/CD systems or source-control platforms, such as Jenkins to GitHub Actions or GitLab to GitHub
- Networking depth: VPC, DNS, load balancers, and WAF
- Leading or co-leading incident response and postmortems
- Experience on a small platform/DevOps team moving fast on multiple initiatives concurrently
- Building or operating AI/LLM backends on AWS, including Bedrock, agentic/tool-calling services, MCP servers, and vector stores
- Hands-on experience with Azure infrastructure
Qualifications
Must Haves
- Bachelor's degree in Computer Science, Engineering, or a related field with 5+ years of relevant experience, or an equivalent combination of education and 8+ years of professional experience in platform engineering, infrastructure, or DevOps
- 5+ years in platform engineering, infrastructure, or DevOps roles, with a proven track record of building internal tooling and Production support
- Strong hands-on experience across AWS infrastructure
- Advanced Terraform: composing and versioning reusable modules; remote state with locking and workspaces/backends; meta-arguments and dynamic blocks (for_each, count, dynamic); safe state operations and refactoring (import, moved blocks, targeted state changes); drift detection; and policy/testing in CI (e.g., Terratest, tflint, Sentinel or OPA)
- Building and debugging CI/CD pipelines independently (GitHub Actions and/or Jenkins)
- Container orchestration on EKS/Kubernetes and AWS Fargate
- GitOps workflows (e.g., ArgoCD)
- Observability in practice: instrumenting services and building meaningful dashboards and alerts across metrics, logs, and APM
- Security fundamentals applied by default: least-privilege IAM, secrets handling, key and certificate rotation
- Strong proficiency with python and shell scripting to automate recurring work and build internal tooling
- Clear written and verbal communication — able to plan a piece of work, document the reasoning, and align with the team before building
- AWS certification, such as Solutions Architect Associate/Professional or SysOps/DevOps Engineer
Nice to Haves
- Ruby or Postgres performance tuning experience
- FinOps or cloud-cost optimization: right-sizing, cost tagging, and reporting
- OpenSearch/Elasticsearch operation and migration
- Data warehouse / analytics infrastructure on AWS (Redshift, Glue, Step Functions)
- Disaster recovery and backup/restore testing; exposure to RTO/RPO planning
- Experience migrating between CI/CD systems or source-control platforms, such as Jenkins to GitHub Actions or GitLab to GitHub
- Networking depth: VPC, DNS, load balancers, and WAF
- Leading or co-leading incident response and postmortems
- Experience on a small platform/DevOps team moving fast on multiple initiatives concurrently
- Building or operating AI/LLM backends on AWS, including Bedrock, agentic/tool-calling services, MCP servers, and vector stores
- Hands-on experience with Azure infrastructure