Truthset logo
Truthset
Posted 75 days agoVerified live 2d ago

Data Engineer

Brief overview

Remote
UndergradOr in progress
$150k–$180k/yrStated range
3+ yrsMinimum
PythonScalaJavaSparkAWS EMRSnowflakeDatabricksRedshiftSQLGraph databasesGithubTerraformAirflowDBT

About the company

Truthset logo
Truthsettruthset.io

Truthset is a data intelligence company that focuses exclusively on validating the accuracy and compliance of consumer data.

Job description

Summary

Truthset is a dynamic and innovative data intelligence company that specializes in validating the accuracy of the world's consumer data to empower data-driven decision-making and marketing success. They are seeking a Senior Data Engineer to design, build, and maintain scalable data pipelines, automate data delivery, and collaborate with the data science team to enhance data processing infrastructure.

Responsibilities

  • Design, build, and maintain scalable data pipelines that supply big data to internal and external teams
  • Automate the delivery of terabytes of structured data to a growing group of enterprise clients
  • Automate the ingestion of terabytes of external data sources into internal data warehouses in different environments (e.g., AWS, Snowflake, Databricks)
  • Write, test, debug, and optimize custom Scala code for ETL workflows and other one-off tasks
  • Deploy ETL code in the cloud (using batch orchestration tools, like Airflow)
  • Work closely with the Head of Data Science and Principal ML Engineer to test and deploy new infrastructure for data processing
  • Create an internal toolkit (KPIs, testing programs, dashboards) to monitor the health of data pipelines
  • Maintain documentation about generated datasets (data dictionaries, feed specs. etc.) for internal and external use
  • Advise the Head of Data Science on future tooling upgrades

Skills

  • Bachelor's in Computer Science, Mathematics, Statistics, or other related fields
  • 3+ years of relevant work experience
  • Proficiency in one or more programming languages such as Python, Scala, Java, or other languages commonly used in data engineering
  • Experience with cloud/distributed computing tools, including Spark, AWS EMR, and cloud-based data warehouse platforms such as Snowflake, Databricks or Redshift
  • A strong background in at least one of the following: distributed data processing or software engineering of data services, or data modeling
  • Experience with relational (SQL) databases and graph databases
  • Experience with version control software, such as Github
  • Excellent communication and collaboration skills
  • Strong problem-solving skills and attention to detail
  • Industry experience programming in Scala
  • Familiarity with a scripting language like Python or R
  • Familiarity with Terraform and Airflow
  • Familiarity with DBT

Qualifications

Must Haves

  • Bachelor's in Computer Science, Mathematics, Statistics, or other related fields
  • 3+ years of relevant work experience
  • Proficiency in one or more programming languages such as Python, Scala, Java, or other languages commonly used in data engineering
  • Experience with cloud/distributed computing tools, including Spark, AWS EMR, and cloud-based data warehouse platforms such as Snowflake, Databricks or Redshift
  • A strong background in at least one of the following: distributed data processing or software engineering of data services, or data modeling
  • Experience with relational (SQL) databases and graph databases
  • Experience with version control software, such as Github
  • Excellent communication and collaboration skills
  • Strong problem-solving skills and attention to detail

Nice to Haves

  • Industry experience programming in Scala
  • Familiarity with a scripting language like Python or R
  • Familiarity with Terraform and Airflow
  • Familiarity with DBT

Benefits

  • Full health benefits
  • 401k
  • The potential for an equity stake

More jobs like this