Summary
GLOBO is a B2B communication platform provider, specializing in translation and interpretation technology, services, data, and insights. The AI/ML Engineer is responsible for building and maintaining data pipelines, AI models, and intelligent features that enhance the GLOBO platform, ensuring high quality and operational efficiency.
Responsibilities
- Build and maintain reliable ingestion pipelines using Snowflake Openflow, Python, and Snowflake, including API, PostgreSQL, and CDC-based integrations
- Develop incremental synchronization, cursor/state management, retry logic, schema-drift handling, soft-delete propagation, and source-to-target reconciliation
- Transform raw source data through staging, intermediate, and core models into trusted datasets for analytics, reporting, and machine-learning workloads
- Apply data-quality checks for freshness, completeness, uniqueness, referential integrity, valid relationships, and business-rule compliance
- Maintain source definitions, model documentation, lineage, metadata, and data contracts
- Collaborate with data owners to ensure PII/PHI classification, masking, retention, and deletion requirements are implemented throughout the pipeline
- Implement monitoring and alerting for ingestion failures, pipeline freshness, schema changes, data-quality failures, transformation errors, model drift, and inference degradation
- Establish automated regression testing for dbt models, features, evaluation datasets, prompts, and model outputs
- Validate that sensitive data is appropriately masked, redacted, access-controlled, and excluded from unauthorized model training or data-sharing workflows
- Build safeguards for PII/PHI in recorded-call, transcript, and AI/ML processing pipelines, including verification that redaction and deletion workflows complete successfully
- Ensure AI/ML outputs are traceable to their source data, model or prompt version, feature set, and evaluation results
- Define recovery procedures, data-quality escalation paths, and operational runbooks for critical pipelines and models
- Support human review and approval for model outputs that may affect customers, interpreters, employees, financial activity, or service quality
- Collaborate with Product and Engineering to ship AI-powered features into the GLOBO platform
- Build and deploy LLM integrations (AWS Bedrock, Anthropic Claude) and agentic workflows (CrewAI, LangChain)
- Write production-quality code with proper tests, documentation, and error handling
- Implement guardrails, monitoring, and alerting for AI services in production
- Ensure AI outputs are consistent and trustworthy
- Contribute to evaluation datasets, prompt versioning, and regression testing for deployed models
- Monitor and optimize Snowflake compute, storage, query performance, dbt execution, Openflow runtime usage, and model-inference costs
- Design efficient incremental models, CDC pipelines, materializations, clustering strategies, and warehouse/task schedules
- Compare and optimize ingestion costs as GLOBO transitions from Fivetran to Snowflake Openflow
- Reduce unnecessary full refreshes, duplicate processing, excessive data movement, and inefficient feature recomputation
- Optimize model selection, prompt size, token usage, batching, caching, inference frequency, and routing between model providers
- Measure model performance against operational cost, latency, throughput, and data-freshness requirements
- Establish practical service-level targets for critical datasets, transformations, batch jobs, and model-serving workflows
Skills
- Bachelor's Degree in Computer Science, Data Science, Information Systems, or related field
- 2+ years of experience in data engineering, software development, or ML engineering
- Experience with the below tech stack is required: Python (advanced proficiency), SQL (advanced proficiency), LLM Integration (AWS Bedrock, Anthropic Claude, or OpenAI API), dbt (data transformation and testing), Snowflake (or similar cloud data warehouse), AWS Lambda / Serverless architecture
- Fivetran (or similar ELT/ingestion tooling)
- Agentic Frameworks (CrewAI, LangChain, or similar)
- Airflow (or similar workflow orchestration)
- Vector Databases (Pinecone, PGVector, or OpenSearch)
- AWS ECS/EKS
- CDK and CloudFormation for automated deployments
- Ruby on Rails (ability to read/debug core platform code)
- Redis
- PostgreSQL
- React
- Familiarity with model evaluation techniques, prompt engineering, and AI safety best practices
- Experience with Google Docs and Apple/Mac Operating System preferred
- Ability to work independently in a decentralized environment without the reliance on direct authority
- Highest level of personal and professional integrity and ethics
- Broad understanding of current and emerging technology practices
- High level of initiative, accountability, and follow-through
- Value strong teamwork and collaboration skills
- Demonstrated problem-solving and decision-making skills
- Ability to manage multiple initiatives and projects and prioritize needs
- Strong sense of service and passion for the company and business
- Authorized to legally work for any employer in the United States
- Willingness to submit to any requested background checks
- Fluent in English
Qualifications
Must Haves
- Bachelor's Degree in Computer Science, Data Science, Information Systems, or related field
- 2+ years of experience in data engineering, software development, or ML engineering
- Experience with the below tech stack is required: Python (advanced proficiency), SQL (advanced proficiency), LLM Integration (AWS Bedrock, Anthropic Claude, or OpenAI API), dbt (data transformation and testing), Snowflake (or similar cloud data warehouse), AWS Lambda / Serverless architecture
Nice to Haves
- Fivetran (or similar ELT/ingestion tooling)
- Agentic Frameworks (CrewAI, LangChain, or similar)
- Airflow (or similar workflow orchestration)
- Vector Databases (Pinecone, PGVector, or OpenSearch)
- AWS ECS/EKS
- CDK and CloudFormation for automated deployments
- Ruby on Rails (ability to read/debug core platform code)
- Redis
- PostgreSQL
- React
- Familiarity with model evaluation techniques, prompt engineering, and AI safety best practices
- Experience with Google Docs and Apple/Mac Operating System preferred
- Ability to work independently in a decentralized environment without the reliance on direct authority
- Highest level of personal and professional integrity and ethics
- Broad understanding of current and emerging technology practices
- High level of initiative, accountability, and follow-through
- Value strong teamwork and collaboration skills
- Demonstrated problem-solving and decision-making skills
- Ability to manage multiple initiatives and projects and prioritize needs
- Strong sense of service and passion for the company and business
- Authorized to legally work for any employer in the United States
- Willingness to submit to any requested background checks
- Fluent in English