Summary
Function Health is building an AI-powered health platform that provides comprehensive, continuous insight into human biology. The Data Engineer will build platform infrastructure that enables engineers, data scientists, clinicians, and agents to safely manage data and ML systems through contracts, testing, observability, lineage, and rollback capabilities.
Responsibilities
- **Tracking infrastructure.** The event pipeline behind product analytics, experimentation, and feature gates. Schemas that are enforced at the source, so a bad event never becomes a bad metric
- **Data processing infrastructure.** The Bronze → Silver → Gold layer in Databricks. Automated schema evolution, contract tests, backfills that aren't scary, freshness and volume monitors generated from the contract rather than bolted on after
- **ML infrastructure.** Feature computation and serving, training and eval pipelines, and the plumbing that gets model output back into the product — with the same testing and observability bar as everything else
- **Cutting across all three.** The self-service story. Templates, local dev and preview environments, policy-as-code for PHI, ownership routing for alerts, and progressive gates so an exploratory model ships freely while a member-facing one earns more scrutiny
Skills
- Built internal platform or infrastructure that other engineers actually adopted
- Operated production data or ML systems, on call for them, and fixed them under pressure
- Strong Python and SQL. Comfortable in a lakehouse – we use Databricks; Snowflake or BigQuery translates fine
- Designed interfaces and schemas that other teams depend on, then evolved them without breaking those teams
- Thought hard about testing and CI for data or ML, where correctness is statistical and the failure is often silent
- 1-4 years of engineering experience gets you here, but we care about what you've built, not the number
- Declarative pipeline frameworks (dbt, DLT, Dagster, Airflow)
- Streaming (Kafka, Spark Structured Streaming)
- Data contracts, data diffing, or lineage tooling
- Terraform
- Feature stores
- MLOps and eval tooling
- Agentic coding workflows
- Healthcare
- PHI, or HIPAA experience
Qualifications
Must Haves
- Built internal platform or infrastructure that other engineers actually adopted
- Operated production data or ML systems, on call for them, and fixed them under pressure
- Strong Python and SQL. Comfortable in a lakehouse – we use Databricks; Snowflake or BigQuery translates fine
- Designed interfaces and schemas that other teams depend on, then evolved them without breaking those teams
- Thought hard about testing and CI for data or ML, where correctness is statistical and the failure is often silent
- 1-4 years of engineering experience gets you here, but we care about what you've built, not the number
Nice to Haves
- Declarative pipeline frameworks (dbt, DLT, Dagster, Airflow)
- Streaming (Kafka, Spark Structured Streaming)
- Data contracts, data diffing, or lineage tooling
- Terraform
- Feature stores
- MLOps and eval tooling
- Agentic coding workflows
- Healthcare
- PHI, or HIPAA experience
Benefits