Summary
Truthset is a dynamic and innovative data intelligence company that specializes in validating the accuracy of the world's consumer data to empower data-driven decision-making and marketing success. They are seeking a Senior Data Engineer to design, build, and maintain scalable data pipelines, automate data delivery, and collaborate with the data science team to enhance data processing infrastructure.
Responsibilities
- Design, build, and maintain scalable data pipelines that supply big data to internal and external teams
- Automate the delivery of terabytes of structured data to a growing group of enterprise clients
- Automate the ingestion of terabytes of external data sources into internal data warehouses in different environments (e.g., AWS, Snowflake, Databricks)
- Write, test, debug, and optimize custom Scala code for ETL workflows and other one-off tasks
- Deploy ETL code in the cloud (using batch orchestration tools, like Airflow)
- Work closely with the Head of Data Science and Principal ML Engineer to test and deploy new infrastructure for data processing
- Create an internal toolkit (KPIs, testing programs, dashboards) to monitor the health of data pipelines
- Maintain documentation about generated datasets (data dictionaries, feed specs. etc.) for internal and external use
- Advise the Head of Data Science on future tooling upgrades
Skills
- Bachelor's in Computer Science, Mathematics, Statistics, or other related fields
- 3+ years of relevant work experience
- Proficiency in one or more programming languages such as Python, Scala, Java, or other languages commonly used in data engineering
- Experience with cloud/distributed computing tools, including Spark, AWS EMR, and cloud-based data warehouse platforms such as Snowflake, Databricks or Redshift
- A strong background in at least one of the following: distributed data processing or software engineering of data services, or data modeling
- Experience with relational (SQL) databases and graph databases
- Experience with version control software, such as Github
- Excellent communication and collaboration skills
- Strong problem-solving skills and attention to detail
- Industry experience programming in Scala
- Familiarity with a scripting language like Python or R
- Familiarity with Terraform and Airflow
- Familiarity with DBT
Qualifications
Must Haves
- Bachelor's in Computer Science, Mathematics, Statistics, or other related fields
- 3+ years of relevant work experience
- Proficiency in one or more programming languages such as Python, Scala, Java, or other languages commonly used in data engineering
- Experience with cloud/distributed computing tools, including Spark, AWS EMR, and cloud-based data warehouse platforms such as Snowflake, Databricks or Redshift
- A strong background in at least one of the following: distributed data processing or software engineering of data services, or data modeling
- Experience with relational (SQL) databases and graph databases
- Experience with version control software, such as Github
- Excellent communication and collaboration skills
- Strong problem-solving skills and attention to detail
Nice to Haves
- Industry experience programming in Scala
- Familiarity with a scripting language like Python or R
- Familiarity with Terraform and Airflow
- Familiarity with DBT
Benefits
- Full health benefits
- 401k
- The potential for an equity stake