Empowers Staffing Inc logo
Empowers Staffing Inc
Posted 81 days agoVerified live 2d ago

AI Data Engineer (ML Data Pipelines)

Brief overview

Remote
4+ yrsMinimum
PythonSQLSparkDatabricksAirflowData PipelinesData Quality ValidationGreat ExpectationsAWSAzureGCPKafkaFeature EngineeringMachine LearningVersion Control

About the company

Empowers Staffing Inc logo
Empowers Staffing Incempowersstaffing.com

Welcome to Empowers Staffing – Where Expertise Meets Diversity! We are more than just a staffing company; we are pioneers in the realm of IT, Banking & Finance Recruitment and Staffing, and we’re growing faster than ever before.

Job description

Summary

Empowers Staffing Inc is seeking an AI Data Engineer to design and build production-grade data pipelines that power machine learning systems. This role focuses on creating scalable ingestion, transformation, and feature engineering workflows that support model training, evaluation, and real-time inference.

Responsibilities

  • Design and build scalable data pipelines for ML workflows
  • Develop feature engineering and data preparation processes
  • Implement batch and real-time data ingestion systems
  • Ensure data quality, validation, and monitoring
  • Collaborate with ML engineers to support model training and deployment
  • Integrate pipelines with orchestration tools (Airflow or similar)
  • Optimize pipeline performance and cloud cost efficiency
  • Maintain documentation and version control of data workflows

Skills

  • 4+ years of experience in Data Engineering
  • Strong Python and SQL skills
  • Experience building data pipelines for ML or analytics systems
  • Hands-on experience with Spark, Databricks, or similar distributed processing frameworks
  • Experience with orchestration tools (Airflow or similar)
  • Experience in AWS, Azure, or GCP environments
  • Familiarity with data quality validation and monitoring frameworks
  • Understanding of feature engineering and model data lifecycle
  • Experience with streaming systems (Kafka, Kinesis, Pub/Sub)
  • Experience supporting model deployment and MLOps workflows
  • Experience with feature stores or vector databases
  • Familiarity with ML frameworks (TensorFlow, PyTorch)

Qualifications

Must Haves

  • 4+ years of experience in Data Engineering
  • Strong Python and SQL skills
  • Experience building data pipelines for ML or analytics systems
  • Hands-on experience with Spark, Databricks, or similar distributed processing frameworks
  • Experience with orchestration tools (Airflow or similar)
  • Experience in AWS, Azure, or GCP environments
  • Familiarity with data quality validation and monitoring frameworks
  • Understanding of feature engineering and model data lifecycle

Nice to Haves

  • Experience with streaming systems (Kafka, Kinesis, Pub/Sub)
  • Experience supporting model deployment and MLOps workflows
  • Experience with feature stores or vector databases
  • Familiarity with ML frameworks (TensorFlow, PyTorch)

More jobs like this