DetailPage logo
DetailPage
Posted 25 days agoVerified live 1d ago

Data Engineer II

Brief overview

Remote
UndergradOr in progress
2+ yrsMinimum
SQLPythonAWSETL/ELTLinuxDockerGitPyTestFastAPIPandasInfrastructure as Code (CDK/Terraform)Aurora PostgreSQL

About the company

DetailPage logo
DetailPagedetailpage.com

DetailPage provides software for optimizing product content and shopping discovery across ecommerce marketplaces and AI tools.

Job description

Summary

DetailPage.com helps Amazon brands grow organic traffic through AI-assisted content optimization and market segment share insights. The Data Engineer II will build and optimize scalable AWS data pipelines, ETL/ELT processes, and API integrations while ensuring reliable data access across the organization and to end users. The role collaborates with engineering, analytics, and product teams and supports testing, infrastructure as code, and CI/CD workflows.

Responsibilities

  • Design, develop, and maintain scalable data pipelines using a variety of AWS services
  • Work with relational databases such as Aurora (PostgreSQL), Redshift, and AWS Glue for large-scale data management
  • Develop and optimize ETL/ELT processes for smooth data integration from multiple sources
  • Collaborate with cross-functional teams to ensure data accessibility and integrity across platforms
  • Use Python, SQL, and Pandas for data manipulation, analysis, and workflow automation
  • Implement unit testing (PyTest) to ensure code reliability
  • Manage API integrations and development using FastAPI
  • Assist with infrastructure as code (IaC) using CDK and Terraform
  • Build CI/CD pipelines in Jenkins

Skills

  • • SQL mastery – expertise in complex SQL queries and database optimization
  • • 2-3 years of experience in Data Engineering with AWS and Python
  • • Linux proficiency – comfortable with shell scripting, common CLI tools, and building/testing Linux-based Docker images
  • • Advanced Python – strong hands-on experience, ideally with Data Engineering tools like PySpark, Pandas, and SQLAlchemy
  • • Git proficiency – comfortable with version control, branching, and collaboration
  • • ETL/ELT experience – proven ability to build and optimize extract, transform, and load processes
  • • Unit testing – familiarity with PyTest or similar tools
  • • AWS experience with Aurora (PostgreSQL or MySQL), Redshift, Fargate, Lambda, SQS, and IAM, plus tools like Boto3 and the AWS CLI
  • • Airflow expertise, particularly AWS Managed Airflow, for scheduling and orchestrating data workflows
  • • API development experience with API Gateway and FastAPI
  • • Infrastructure as Code experience with CDK or Terraform
  • • Familiarity with caching tools like Redis
  • • Basic DevOps knowledge for CI/CD pipelines and dev workflow optimization
  • • Familiarity with DBT

Qualifications

Must Haves

  • • SQL mastery – expertise in complex SQL queries and database optimization
  • • 2-3 years of experience in Data Engineering with AWS and Python
  • • Linux proficiency – comfortable with shell scripting, common CLI tools, and building/testing Linux-based Docker images
  • • Advanced Python – strong hands-on experience, ideally with Data Engineering tools like PySpark, Pandas, and SQLAlchemy
  • • Git proficiency – comfortable with version control, branching, and collaboration
  • • ETL/ELT experience – proven ability to build and optimize extract, transform, and load processes
  • • Unit testing – familiarity with PyTest or similar tools

Nice to Haves

  • • AWS experience with Aurora (PostgreSQL or MySQL), Redshift, Fargate, Lambda, SQS, and IAM, plus tools like Boto3 and the AWS CLI
  • • Airflow expertise, particularly AWS Managed Airflow, for scheduling and orchestrating data workflows
  • • API development experience with API Gateway and FastAPI
  • • Infrastructure as Code experience with CDK or Terraform
  • • Familiarity with caching tools like Redis
  • • Basic DevOps knowledge for CI/CD pipelines and dev workflow optimization
  • • Familiarity with DBT

More jobs like this