Summary
Dice is seeking an experienced Data Engineer to join its team. The role focuses on building robust, scalable PySpark data pipelines, optimizing Spark workloads, maintaining ETL/ELT processes, resolving data issues, and collaborating with data teams to ensure data quality and reliability.
Responsibilities
- Design and implement scalable PySpark data pipelines for batch and streaming workloads
- Optimize Spark jobs and queries for performance and cost efficiency
- Build and maintain ETL/ELT processes following data engineering best practices
- Troubleshoot and resolve complex data pipeline and processing issues
- Collaborate with data teams to ensure data quality and reliability
Skills
- 3+ years of hands-on experience building data pipelines in Databricks; deep understanding of Spark fundamentals, transformations, actions, and performance optimization techniques including partitioning, caching, and resource management
- Expert-level proficiency writing production-quality PySpark code and complex SQL queries for data transformation, aggregation, and analysis; experience with DataFrame API, Spark SQL, and UDFs; strong understanding of lazy evaluation and execution plans
- Proven experience building and maintaining production data pipelines; hands-on experience with incremental data loading, change data capture (CDC), and slowly changing dimensions; experience handling data quality issues and implementing data validation frameworks
- Strong proficiency with AWS services (S3, EC2, IAM, Glue, Athena); experience working with large-scale distributed data processing; familiarity with data formats (Parquet, Delta, JSON, Avro) and compression techniques
- Experience with version control (Git) and CI/CD pipelines using GitLab, GitHub Actions, or similar tools; familiarity with testing data pipelines and deployment automation; experience with Databricks Repos and workspace-level integrations
- Understanding of data lineage, cataloging, and metadata management; experience implementing data quality checks and monitoring; knowledge of data privacy and security best practices in cloud environments (nice to have)
- Knowledge of data privacy and security best practices in cloud environments (nice to have)
Qualifications
Must Haves
- 3+ years of hands-on experience building data pipelines in Databricks; deep understanding of Spark fundamentals, transformations, actions, and performance optimization techniques including partitioning, caching, and resource management
- Expert-level proficiency writing production-quality PySpark code and complex SQL queries for data transformation, aggregation, and analysis; experience with DataFrame API, Spark SQL, and UDFs; strong understanding of lazy evaluation and execution plans
- Proven experience building and maintaining production data pipelines; hands-on experience with incremental data loading, change data capture (CDC), and slowly changing dimensions; experience handling data quality issues and implementing data validation frameworks
- Strong proficiency with AWS services (S3, EC2, IAM, Glue, Athena); experience working with large-scale distributed data processing; familiarity with data formats (Parquet, Delta, JSON, Avro) and compression techniques
- Experience with version control (Git) and CI/CD pipelines using GitLab, GitHub Actions, or similar tools; familiarity with testing data pipelines and deployment automation; experience with Databricks Repos and workspace-level integrations
- Understanding of data lineage, cataloging, and metadata management; experience implementing data quality checks and monitoring; knowledge of data privacy and security best practices in cloud environments (nice to have)
Nice to Haves
- Knowledge of data privacy and security best practices in cloud environments (nice to have)
Benefits