Summary
DetailPage.com helps Amazon brands grow organic traffic through AI-assisted content optimization and market segment share insights. The Data Engineer II will build and optimize scalable AWS data pipelines, ETL/ELT processes, and API integrations while ensuring reliable data access across the organization and to end users. The role collaborates with engineering, analytics, and product teams and supports testing, infrastructure as code, and CI/CD workflows.
Responsibilities
- Design, develop, and maintain scalable data pipelines using a variety of AWS services
- Work with relational databases such as Aurora (PostgreSQL), Redshift, and AWS Glue for large-scale data management
- Develop and optimize ETL/ELT processes for smooth data integration from multiple sources
- Collaborate with cross-functional teams to ensure data accessibility and integrity across platforms
- Use Python, SQL, and Pandas for data manipulation, analysis, and workflow automation
- Implement unit testing (PyTest) to ensure code reliability
- Manage API integrations and development using FastAPI
- Assist with infrastructure as code (IaC) using CDK and Terraform
- Build CI/CD pipelines in Jenkins
Skills
- • SQL mastery – expertise in complex SQL queries and database optimization
- • 2-3 years of experience in Data Engineering with AWS and Python
- • Linux proficiency – comfortable with shell scripting, common CLI tools, and building/testing Linux-based Docker images
- • Advanced Python – strong hands-on experience, ideally with Data Engineering tools like PySpark, Pandas, and SQLAlchemy
- • Git proficiency – comfortable with version control, branching, and collaboration
- • ETL/ELT experience – proven ability to build and optimize extract, transform, and load processes
- • Unit testing – familiarity with PyTest or similar tools
- • AWS experience with Aurora (PostgreSQL or MySQL), Redshift, Fargate, Lambda, SQS, and IAM, plus tools like Boto3 and the AWS CLI
- • Airflow expertise, particularly AWS Managed Airflow, for scheduling and orchestrating data workflows
- • API development experience with API Gateway and FastAPI
- • Infrastructure as Code experience with CDK or Terraform
- • Familiarity with caching tools like Redis
- • Basic DevOps knowledge for CI/CD pipelines and dev workflow optimization
- • Familiarity with DBT
Qualifications
Must Haves
- • SQL mastery – expertise in complex SQL queries and database optimization
- • 2-3 years of experience in Data Engineering with AWS and Python
- • Linux proficiency – comfortable with shell scripting, common CLI tools, and building/testing Linux-based Docker images
- • Advanced Python – strong hands-on experience, ideally with Data Engineering tools like PySpark, Pandas, and SQLAlchemy
- • Git proficiency – comfortable with version control, branching, and collaboration
- • ETL/ELT experience – proven ability to build and optimize extract, transform, and load processes
- • Unit testing – familiarity with PyTest or similar tools
Nice to Haves
- • AWS experience with Aurora (PostgreSQL or MySQL), Redshift, Fargate, Lambda, SQS, and IAM, plus tools like Boto3 and the AWS CLI
- • Airflow expertise, particularly AWS Managed Airflow, for scheduling and orchestrating data workflows
- • API development experience with API Gateway and FastAPI
- • Infrastructure as Code experience with CDK or Terraform
- • Familiarity with caching tools like Redis
- • Basic DevOps knowledge for CI/CD pipelines and dev workflow optimization
- • Familiarity with DBT