Summary
GLOBO is seeking a Data & ML Engineer to build, operate, and secure the data pipelines and machine-learning infrastructure supporting its analytics, reporting, and AI capabilities. The role is responsible for data ingestion, analytics engineering, data quality and privacy, ML operations, AI enablement, monitoring, and performance optimization across GLOBO's data and AI technology stack.
Responsibilities
- Build and maintain reliable ingestion pipelines using Snowflake Openflow, Fivetran, and Python, including PostgreSQL CDC (logical replication/WAL), API, and file-based integrations
- Implement incremental synchronization, cursor/state management, retry logic, schema-drift handling, soft-delete propagation, and source-to-target reconciliation (row counts, inserts, updates, deletes, lag)
- Monitor CDC health, including replication-slot status, WAL accumulation, connector lag, and failure recovery
- Support the transition of in-scope sources from Fivetran to Snowflake Openflow, including cadence-tier design and cost/latency validation
- Onboard new sources into Snowflake (e.g., operational spreadsheets, SaaS and translation-management APIs) with documented ownership and freshness expectations
- Transform raw source data through staging, intermediate, and mart models in dbt into trusted datasets for analytics, reporting, embedded dashboards, and ML workloads
- Design dimensional and fact models at the correct grain for both batch and near-real-time use cases
- Codify approved business definitions into version-controlled dbt models and support their exposure through the Omni semantic layer
- Maintain source definitions, model documentation, lineage, metadata, and data contracts between Data Engineering and BI
- Investigate and resolve data mismatches between source systems, Snowflake, and BI tools
- Implement data-quality checks for freshness, completeness, uniqueness, referential integrity, and business-rule compliance, and use them as deployment gates
- Implement monitoring and alerting for ingestion failures, pipeline freshness, schema changes, transformation errors, and Snowflake task failures, integrated with team alerting channels and log-monitoring tools
- Own the recorded-call and transcript redaction pipeline (e.g., S3 → transcription → Snowflake → redaction → PostgreSQL/Snowflake), including verification that PII/PHI redaction and deletion workflows complete successfully
- Collaborate with data owners and Security/Compliance to implement PII/PHI classification, masking, row-level access, retention, and deletion requirements throughout the pipeline
- Define recovery procedures, escalation paths, and operational runbooks for critical pipelines and models
- Contribute to CI/CD, branch/approval policy, and infrastructure-as-code for the data platform
- Migrate and operationalize existing ML models (e.g., the staffing/forecasting model) onto Snowflake with MLflow-based experiment tracking, model registry, and versioned deployment
- Build and maintain the inventory of ML and LLM models used across GLOBO applications and services, including owners, inputs, versions, and evaluation status
- Implement recurring drift, performance, and inference-degradation monitoring for production models
- Support evaluation and backtesting workflows for AI features (e.g., Kai), including evaluation datasets, prompt/model versioning, and regression tests
- Build LLM-assisted data workflows where they add clear value (e.g., transcript processing, redaction assistance, metadata/RAG enrichment) using AWS Bedrock/Anthropic Claude, with guardrails and human review
- Make sure ML and AI outputs can be traced to their source data, feature set, model or prompt version, and evaluation results
- Validate that sensitive data is excluded from unauthorized model training or data-sharing workflows
- Monitor and optimize Snowflake compute, storage, and query performance; dbt execution; Openflow runtime usage; and model-inference costs
- Design efficient incremental models, materializations, clustering strategies, and warehouse/task schedules
- Reduce unnecessary full refreshes, duplicate processing, and inefficient feature recomputation
- Establish practical service-level targets for data freshness, transformation latency, and model-serving workflows, and measure against them
Skills
- Bachelor's Degree in Computer Science, Data Science, Information Systems, Engineering, or related field, or equivalent practical experience
- 4+ years of experience in data engineering, analytics engineering, or ML engineering, including ownership of production pipelines
- Python (advanced proficiency)
- SQL (advanced proficiency), including dimensional/data modeling
- Snowflake (or similar cloud data warehouse)
- Dbt (data transformation, testing, and documentation)
- PostgreSQL, including change data capture / logical replication
- ELT/ingestion tooling (Snowflake Openflow, Fivetran, or similar)
- Data-quality testing and pipeline monitoring/observability
- Handling regulated or sensitive data (PII/PHI masking, access control, retention)
- Familiarity with model evaluation techniques, prompt engineering, and responsible-AI practices
- Ability to work independently in a decentralized environment without the reliance on direct authority
- Highest level of personal and professional integrity and ethics
- Broad understanding of current and emerging technology practices
- High level of initiative, accountability, and follow-through
- Value strong teamwork and collaboration skills
- Demonstrated problem-solving and decision-making skills
- Ability to manage multiple initiatives and projects and prioritize needs
- Ability to translate business questions into data models and metrics
- Strong sense of service and passion for the company and business
- Authorized to legally work for any employer in the United States
- Fluent in English
- Drive Results - Take accountability for individual outcomes - good or bad; Prioritize work to support Company rocks
- Communicate Effectively - Listen with intent to understand; Ask questions to ensure understanding
- Developing Self and Others - expand self-awareness and be open to feedback; take ownership/seek opportunities for learning, career growth, and development
- Growth-Focused - Develop a growth mindset by keeping an open mind during change; Seek opportunities to move out of your comfort zone
- Customer Centricity - Consistently provides excellent customer service; Actively gathers and leverages information to understand current and emerging customer priorities, problems, expectations, and needs - presents proposed solutions within areas of responsibility
- MLflow, Snowflake ML/Snowpark, or similar ML lifecycle tooling
- Model monitoring and drift detection
- LLM integration (AWS Bedrock, Anthropic Claude, or OpenAI API)
- Agentic or RAG frameworks (LangChain, CrewAI, or similar) and vector databases (PGVector, OpenSearch, or similar)
- BI/semantic-layer tools (Omni, Looker, or similar)
- CI/CD for data (CircleCI, GitHub Actions, or similar)
- Infrastructure as code (Terraform, CDK, or CloudFormation)
- AWS services (S3, Lambda, Transcribe, VPC/PrivateLink)
- Workflow orchestration (Airflow or Snowflake Tasks)
- Log monitoring tools (Sumo Logic or similar)
- Experience in HIPAA-regulated environments
- Experience with Google Docs and Apple/Mac Operating System
Qualifications
Must Haves
- Bachelor's Degree in Computer Science, Data Science, Information Systems, Engineering, or related field, or equivalent practical experience
- 4+ years of experience in data engineering, analytics engineering, or ML engineering, including ownership of production pipelines
- Python (advanced proficiency)
- SQL (advanced proficiency), including dimensional/data modeling
- Snowflake (or similar cloud data warehouse)
- dbt (data transformation, testing, and documentation)
- PostgreSQL, including change data capture / logical replication
- ELT/ingestion tooling (Snowflake Openflow, Fivetran, or similar)
- Data-quality testing and pipeline monitoring/observability
- Handling regulated or sensitive data (PII/PHI masking, access control, retention)
- Familiarity with model evaluation techniques, prompt engineering, and responsible-AI practices
- Ability to work independently in a decentralized environment without the reliance on direct authority
- Highest level of personal and professional integrity and ethics
- Broad understanding of current and emerging technology practices
- High level of initiative, accountability, and follow-through
- Value strong teamwork and collaboration skills
- Demonstrated problem-solving and decision-making skills
- Ability to manage multiple initiatives and projects and prioritize needs
- Ability to translate business questions into data models and metrics
- Strong sense of service and passion for the company and business
- Authorized to legally work for any employer in the United States
- Fluent in English
- Drive Results - Take accountability for individual outcomes - good or bad; Prioritize work to support Company rocks
- Communicate Effectively - Listen with intent to understand; Ask questions to ensure understanding
- Developing Self and Others - expand self-awareness and be open to feedback; take ownership/seek opportunities for learning, career growth, and development
- Growth-Focused - Develop a growth mindset by keeping an open mind during change; Seek opportunities to move out of your comfort zone
- Customer Centricity - Consistently provides excellent customer service; Actively gathers and leverages information to understand current and emerging customer priorities, problems, expectations, and needs - presents proposed solutions within areas of responsibility
Nice to Haves
- MLflow, Snowflake ML/Snowpark, or similar ML lifecycle tooling
- Model monitoring and drift detection
- LLM integration (AWS Bedrock, Anthropic Claude, or OpenAI API)
- Agentic or RAG frameworks (LangChain, CrewAI, or similar) and vector databases (PGVector, OpenSearch, or similar)
- BI/semantic-layer tools (Omni, Looker, or similar)
- CI/CD for data (CircleCI, GitHub Actions, or similar)
- Infrastructure as code (Terraform, CDK, or CloudFormation)
- AWS services (S3, Lambda, Transcribe, VPC/PrivateLink)
- Workflow orchestration (Airflow or Snowflake Tasks)
- Log monitoring tools (Sumo Logic or similar)
- Experience in HIPAA-regulated environments
- Experience with Google Docs and Apple/Mac Operating System
Benefits
- Hybrid and/or remote work environment