Summary
Alt is unlocking the value of alternative assets through a platform for buying, selling, vaulting, and financing trading cards. The Data Engineer will own end-to-end data pipelines that scrape, normalize, monitor, and ingest transaction and listing data to power Alt Value, pricing, analytics, and downstream products.
Responsibilities
- Design, optimize, and own the pipelines that scrape, process, and ingest transaction and listing data from major auction houses and marketplaces. Scraper to orchestrator to warehouse, end to end
- Build the monitoring and alerting that tracks latency, uptime, and coverage across every source. A source going quiet should page you, not surprise a pricer three days later
- Modernize storage and processing, cut manual intervention out of the loop, and optimize for cost, performance, and reliability. Incremental and evidence-based, not a rewrite
- Partner with pricing, ML, product, and analytics to understand how the data actually gets used, and deliver it clean and standardized enough that nobody downstream is writing defensive code around your output
- Point an agent at a site that just changed its markup and have it propose the parser fix, run it against a fixture set, and open the PR
- Build a triage loop over the orchestrator — cross-reference a failed DAG against logs and warehouse state, diagnose the cause, and hand you a ranked list instead of a wall of red
- Use LLM extraction for the long tail of low-volume sources where a bespoke scraper will never pay for itself
- Put an MCP server over the warehouse so a coverage question is a sentence, not another one-off query
- Use Claude Code or Cursor as the default way you work through a migration or a backfill, not as autocomplete
Skills
- 3-4 years in data engineering or a closely adjacent seat
- Strong Python, 3+ years hands-on, and real large-scale processing with dataframe technologies (Pandas, Polars, PySpark, or similar)
- Orchestration in your hands, not your resume. Airflow, Dagster, or a comparable DAG system — you've built on it, been paged for it, and cleaned up after it
- A pipeline you owned end to end in the past two years. Not a pipeline you contributed to. One that was yours when it broke
- Solid SQL for analysis and transformation
- Startup experience. You understand the pace, and you've worked somewhere the roadmap changed under you
- A pragmatic bias. Value delivered incrementally over the perfect rebuild
- Bonus: web scraping at scale (Selenium, Puppeteer, Beautiful Soup). AWS. LLM-based extraction and processing. You collect, or you have opinions about the collectibles market
Qualifications
Must Haves
- 3-4 years in data engineering or a closely adjacent seat
- Strong Python, 3+ years hands-on, and real large-scale processing with dataframe technologies (Pandas, Polars, PySpark, or similar)
- Orchestration in your hands, not your resume. Airflow, Dagster, or a comparable DAG system — you've built on it, been paged for it, and cleaned up after it
- A pipeline you owned end to end in the past two years. Not a pipeline you contributed to. One that was yours when it broke
- Solid SQL for analysis and transformation
- Startup experience. You understand the pace, and you've worked somewhere the roadmap changed under you
- A pragmatic bias. Value delivered incrementally over the perfect rebuild
Nice to Haves
- Bonus: web scraping at scale (Selenium, Puppeteer, Beautiful Soup). AWS. LLM-based extraction and processing. You collect, or you have opinions about the collectibles market
Benefits
- We cover 85% of your medical, dental, and vision, and up to 50% for dependents. HSA/FSA available
- Generous PTO that people actually take
- Parental leave at full salary
- $200/month wellness stipend
- $100/month home office stipend
- Free DoorDash Pass membership
- Remote-first (for most positions), with WeWork access when you want a room with other humans
- 401(k)
- Alt Equity to all full-time employees