Catena Clearing logo
Catena Clearing
Posted 57 days agoVerified live 13h ago

Data Software Engineer

Brief overview

Remote
3+ yrsMinimum
PythonSQLRelational data modelingStreaming ingestionEvent-driven ingestionAWSAWS S3AWS LambdaAWS KinesisPostgreSQLAWS AuroraSnowflakeWebhooksREST APIsClear written communication

About the company

Catena Clearing logo
Catena Clearingcatenaclearing.io

The Plaid for Trucking

Job description

Summary

Catena Clearing is a startup focused on providing a universal API for fleet telematics data. They are seeking a Data Software Engineer to build and maintain data pipelines that integrate various telematics providers into a coherent data product, ensuring high data quality and operational efficiency.

Responsibilities

  • Build and operate ingestion pipelines across our telematics provider integrations, streaming, polling, and batch - with the retry, rate-limiting, and backpressure behavior each provider demands
  • Own the normalization layer: map wildly inconsistent provider payloads into Catena's canonical vehicle, driver, HOS, and fuel models, and defend that schema as new providers are added
  • Design storage and query patterns for high-volume time-series vehicle data, balancing real-time API latency against analytical workloads
  • Ship customer-facing data products, stop and geofence detection, ETA computation, location aggregates, IFTA-grade mileage - from spec through production
  • Build and maintain write-back workflows that push fuel transactions, DVIR status, and dispatch data back into provider systems, with idempotency and loop protection
  • Instrument everything: data freshness, field completeness, provider health, and drift detection, so we catch a broken upstream before a customer does
  • Own data quality end to end. In our business a silently wrong location is worse than a missing one - a fuel-card issuer declines a real transaction, an insurer misprices a carrier
  • Work directly with customers' engineering teams when the data question is theirs, not ours

Skills

  • 3+ years building production data pipelines or backend data services
  • Strong Python; comfortable owning services end to end, not just notebooks or DAGs
  • Solid SQL and relational data modeling: you can reason about partitioning, indexing, and query cost, not just correctness
  • Experience with streaming or event-driven ingestion, and with the operational reality of unreliable third-party APIs
  • Cloud-native on AWS
  • You've debugged a pipeline that was quietly producing wrong data, and you have opinions about how to prevent it happening again
  • Clear written communication, we're remote and async by default
  • Experience in telematics, IoT, logistics, fintech, or insurance data - anywhere the data has physical-world ground truth and someone makes a money decision on it
  • Time-series or geospatial data at scale: geofencing, map matching, trip and stop inference, H3 or similar indexing
  • Data-sharing and multi-tenant delivery patterns - Snowflake shares, per-tenant credential isolation, consent-scoped access
  • Experience building against many third-party APIs at once, where you control neither the schema nor the uptime
  • Working in a SOC 2 environment, or helping get a company there

Qualifications

Must Haves

  • 3+ years building production data pipelines or backend data services
  • Strong Python; comfortable owning services end to end, not just notebooks or DAGs
  • Solid SQL and relational data modeling: you can reason about partitioning, indexing, and query cost, not just correctness
  • Experience with streaming or event-driven ingestion, and with the operational reality of unreliable third-party APIs
  • Cloud-native on AWS
  • You've debugged a pipeline that was quietly producing wrong data, and you have opinions about how to prevent it happening again
  • Clear written communication, we're remote and async by default

Nice to Haves

  • Experience in telematics, IoT, logistics, fintech, or insurance data - anywhere the data has physical-world ground truth and someone makes a money decision on it
  • Time-series or geospatial data at scale: geofencing, map matching, trip and stop inference, H3 or similar indexing
  • Data-sharing and multi-tenant delivery patterns - Snowflake shares, per-tenant credential isolation, consent-scoped access
  • Experience building against many third-party APIs at once, where you control neither the schema nor the uptime
  • Working in a SOC 2 environment, or helping get a company there

Benefits

  • Full benefits
  • Equity
  • 401k match
  • Remote-friendly environment
  • Occasional travel for team and customer on-sites

More jobs like this