PadSplit logo
PadSplit
Posted 37 days agoVerified live 2d ago

Data Engineer (Fully Remote)

Brief overview

Remote
Dagster or AirflowdbtPythonSnowflakeAWS S3AWS IAMAWS ECS/FargateAirbytePostgreSQLTerraformData WarehousingDimensional Data Modeling

About the company

PadSplit logo
PadSplitpadsplit.com

PadSplit is an affordable housing tech startup that provides a house-sharing service for the workforce.

Job description

Summary

PadSplit is a company focused on solving the affordable housing crisis. The Data Engineer will build and maintain ingestion, transformation, orchestration, and data modeling pipelines across Dagster or Airflow, dbt, Snowflake, Python, Airbyte, and AWS. The role also includes production monitoring, backfills, code reviews, infrastructure updates, and collaboration on data quality and dimensional modeling.

Responsibilities

  • Opening and merging pull requests for new or updated Dagster jobs, assets, schedules, and sensors, plus dbt models, tests, and documentation
  • Writing occasional Terraform for secrets, environment variables, or job sizing when a pipeline needs it
  • Building and debugging Python pipelines covering REST/API syncs, large Postgres extracts, Parquet loads, and Snowflake COPY operations
  • Configuring or troubleshooting Airbyte connections wherever managed sync is the right fit
  • Watching production runs and investigating failures related to IAM, OOM, Spot instances, or bad watermarks
  • Running backfills and incremental catch-ups with a clear story for what landed and why
  • Working with analytics and product on dim/fct/x_fct design, incremental strategies, and data quality
  • Participating in code review, release prep, and writing short runbooks so others can operate your pipelines when you're out

Skills

  • We're looking for a practitioner who thinks natively in dimensions, facts, and slowly changing dimensions — someone who knows when to use a full refresh versus an incremental load and how that choice affects idempotency and backfills
  • This person writes clear, reviewable PRs and gives equally thoughtful reviews, with attention to scoped diffs, sensible tests, and failure modes
  • They're comfortable in complex Python data flows and have enough AWS literacy to reason about task roles, buckets, and cross-account access without needing to own all of platform engineering
  • Solid grasp of relational databases and warehouse patterns — keys, grain, normalization vs. star schema, and how SCD behavior gets encoded
  • Practical, hands-on Dagster (or Airflow) experience — not just writing SQL inside a scheduler UI
  • Real experience building and maintaining models, tests, and documentation in dbt
  • Comfort reading and writing Python that moves data at scale across extract, transform, and load steps
  • Practical familiarity with S3, IAM, and ECS/Fargate at a "debug my job" level
  • Experience with Airbyte or similar extract-and-load tools
  • The discipline to write pull requests others can easily review, and to give equally rigorous reviews in return
  • A track record of keeping pipelines healthy and modeling consistent across full refresh and incremental paths, without becoming a single point of failure

Qualifications

Must Haves

  • We're looking for a practitioner who thinks natively in dimensions, facts, and slowly changing dimensions — someone who knows when to use a full refresh versus an incremental load and how that choice affects idempotency and backfills
  • This person writes clear, reviewable PRs and gives equally thoughtful reviews, with attention to scoped diffs, sensible tests, and failure modes
  • They're comfortable in complex Python data flows and have enough AWS literacy to reason about task roles, buckets, and cross-account access without needing to own all of platform engineering
  • Solid grasp of relational databases and warehouse patterns — keys, grain, normalization vs. star schema, and how SCD behavior gets encoded
  • Practical, hands-on Dagster (or Airflow) experience — not just writing SQL inside a scheduler UI
  • Real experience building and maintaining models, tests, and documentation in dbt
  • Comfort reading and writing Python that moves data at scale across extract, transform, and load steps
  • Practical familiarity with S3, IAM, and ECS/Fargate at a "debug my job" level
  • Experience with Airbyte or similar extract-and-load tools
  • The discipline to write pull requests others can easily review, and to give equally rigorous reviews in return
  • A track record of keeping pipelines healthy and modeling consistent across full refresh and incremental paths, without becoming a single point of failure

Benefits

  • Fully remote position
  • Competitive compensation package including an equity incentive plan and company-wide bonus opportunity
  • National medical, dental, and vision healthcare plans
  • Company provided life insurance policy
  • Optional accidental insurances, FSA, and DCFSA benefits
  • Unlimited paid-time (PTO) policy with eleven (11) company-observed holidays
  • 401(k) plan
  • Twelve (12) weeks of paid time off for both birth and non-birth parents
  • The opportunity to do what you love at a company that is at the forefront of solving the affordable housing crisis

More jobs like this