Summary
UNFI is a premier North American grocery wholesaler delivering fresh, branded, and owned-brand products to community grocers and retail chains. The IT Site Reliability Engineer II operates, secures, optimizes, and continuously improves enterprise data platforms across Databricks, AWS, SAP Business Data Cloud, and legacy environments. The role focuses on platform administration, security, cost optimization, automation, AI-enabled DataOps, monitoring, troubleshooting, and production reliability.
Responsibilities
- Databricks Administration: Operate and maintain Databricks workspaces, jobs, clusters, SQL warehouses, Delta Lake, Unity Catalog, cluster policies, configurations, and environments
- AWS Platform Operations: Support AWS data infrastructure including S3, Glue, Redshift, Lambda, CloudWatch, networking, and storage and other AWS services as required
- IAM & Security: Administer AWS IAM and Databricks access controls, including RBAC/ABAC, permissions, service accounts, and best security practices
- Platform Lifecycle & Hygiene: Manage upgrades, patches, configurations, housekeeping, cleanup, resource lifecycle, and platform standards to maintain a healthy environment
- Multi-Platform Operations: Maintain operational stability and integration across SAP Business Data Cloud/Datasphere, Teradata, SQL Server, Informatica, DataStage, Fivetran, and HVR
- Cloud & Databricks FinOps: Monitor and optimize AWS and Databricks spend, including DBUs, compute, SQL warehouses, storage, and data processing
- Cost Governance: Use AWS Cost Explorer, AWS Budgets, and Databricks billing data to improve cost visibility, establish alerts, and drive optimization opportunities
- Automation & IaC: Automate platform administration, provisioning, configuration, monitoring, and maintenance using Terraform, Python, SQL, APIs, and CI/CD
- Operational Standards: Establish repeatable practices for platform maintenance, monitoring, access management, resource utilization, and operational governance
- Intelligent Operations: Apply AI/LLMs to incident triage, root-cause analysis, anomaly detection, operational monitoring, and platform intelligence
- Agentic DataOps: Develop AI-assisted runbooks, intelligent workflows, automated remediation, and predictive monitoring capabilities
- Proactive Operations: Use AI and automation to move from reactive support toward proactive, predictive, and increasingly self-service DataOps
- Performs other duties as assigned
Skills
- Bachelor's degree in computer science, data analytics, systems analysis, or a related field
- 3-5 years of experience in Platform Engineering, Cloud Operations, DataOps, Production Data Engineering, or related technology operations
- Strong hands-on experience administering Databricks and AWS production environments
- Strong understanding of Databricks administration, Unity Catalog, cluster policies, AWS IAM, security controls, and platform configuration
- Demonstrated experience with cloud/Databricks FinOps and cost optimization
- Strong Python and SQL skills with experience automating operational processes
- Experience with Terraform/IaC, CI/CD, monitoring and ITSM processes
- Demonstrated hands-on experience applying AI/LLMs, intelligent automation, anomaly detection, or AI-assisted operational workflows
- Strong production troubleshooting, incident management, RCA, reliability, and platform maintenance experience
- Familiarity with ingestion tools (Fivetran HVR, AWS DMS, DataStage, Informatica) and BI platforms (Power BI, Tableau, Alteryx)
- Experience with SAP, master data management, and cross-functional processes across supply chain, finance, and operations
- Databricks Certified Platform Administrator
- Databricks Mosaic AI, Genie, AI/BI, agentic workflows, or LLM frameworks
- SAP Business Data Cloud / SAP Datasphere
- Teradata, SQL Server, Informatica, DataStage, Fivetran or HVR
- Enterprise observability platforms such as Datadog, Prometheus, Monte Carlo, or CloudWatch
- Experience operating hybrid cloud and legacy-to-cloud data environments
- Strong troubleshooting and incident management skills
- Knowledge of governance, security, and RBAC principles
- Ability to work independently and collaborate with external partners
- Familiarity with Agile practices and DevOps principles. Understanding of governance, security, and privacy
- Demonstrated commitment to continuous learning, professional development, and staying current with emerging technologies, cloud platforms, AI capabilities, and industry best practices
- Good judgment is required for this position as there may be times when direct supervision may not be immediately available
- This position is classified as remote where the associate will perform remote work from their primary residence
- This position may require the associate to travel to company offices, distribution centers, or other locations for specific meetings or other business reasons
- Incumbent may sit for long periods of time at a desk or computer terminal
- While performing the duties of this job, the employee is regularly required to sit; use hands to finger, handle, or feel; reach with hands and arms; and talk or hear
- Incumbent may use calculators, keyboards, telephones, and other office equipment in the course of a normal workday
- Stooping, bending, twisting, and reaching may be required in the completion of job duties
Qualifications
Must Haves
- Bachelor's degree in computer science, data analytics, systems analysis, or a related field
- 3-5 years of experience in Platform Engineering, Cloud Operations, DataOps, Production Data Engineering, or related technology operations
- Strong hands-on experience administering Databricks and AWS production environments
- Strong understanding of Databricks administration, Unity Catalog, cluster policies, AWS IAM, security controls, and platform configuration
- Demonstrated experience with cloud/Databricks FinOps and cost optimization
- Strong Python and SQL skills with experience automating operational processes
- Experience with Terraform/IaC, CI/CD, monitoring and ITSM processes
- Demonstrated hands-on experience applying AI/LLMs, intelligent automation, anomaly detection, or AI-assisted operational workflows
- Strong production troubleshooting, incident management, RCA, reliability, and platform maintenance experience
- Familiarity with ingestion tools (Fivetran HVR, AWS DMS, DataStage, Informatica) and BI platforms (Power BI, Tableau, Alteryx)
- Experience with SAP, master data management, and cross-functional processes across supply chain, finance, and operations
- Databricks Certified Platform Administrator
- Databricks Mosaic AI, Genie, AI/BI, agentic workflows, or LLM frameworks
- SAP Business Data Cloud / SAP Datasphere
- Teradata, SQL Server, Informatica, DataStage, Fivetran or HVR
- Enterprise observability platforms such as Datadog, Prometheus, Monte Carlo, or CloudWatch
- Experience operating hybrid cloud and legacy-to-cloud data environments
- Strong troubleshooting and incident management skills
- Knowledge of governance, security, and RBAC principles
- Ability to work independently and collaborate with external partners
- Familiarity with Agile practices and DevOps principles. Understanding of governance, security, and privacy
- Demonstrated commitment to continuous learning, professional development, and staying current with emerging technologies, cloud platforms, AI capabilities, and industry best practices
- Good judgment is required for this position as there may be times when direct supervision may not be immediately available
- This position is classified as remote where the associate will perform remote work from their primary residence
- This position may require the associate to travel to company offices, distribution centers, or other locations for specific meetings or other business reasons
- Incumbent may sit for long periods of time at a desk or computer terminal
- While performing the duties of this job, the employee is regularly required to sit; use hands to finger, handle, or feel; reach with hands and arms; and talk or hear
- Incumbent may use calculators, keyboards, telephones, and other office equipment in the course of a normal workday
- Stooping, bending, twisting, and reaching may be required in the completion of job duties
Benefits
- Competitive 401k
- Flexible PTO
- Remote
- Health benefits – first of the month following 30 days of employment
- Mentorship program/developmental opportunities
- Paid Time Off
- Sick Time
- Paid holidays and parental leave
- 401K Program
- Medical, dental, vision, life, and accidental death/dismemberment insurance
- Short-term and long-term disability insurance program
- Flexible Spending Account and/or Health Savings Account, subject to meeting the eligibility requirements and the terms and conditions of these programs, and subject to any requirements under applicable collective bargaining agreements