Summary
b.well Connected Health is seeking a Data Scientist to support Bailey, its clinical AI system. The role owns statistical analysis for drift detection, safety-weighted quality scoring, confidence calibration, FHIR ground-truth comparisons, health scoring, RAG and embedding evaluation, and clinical skills library analysis.
Responsibilities
- Evaluation science support (pairs with the new ML/Platform Engineer): once the platform engineer builds the pipelines, you own the analytical side — defining drift thresholds, the statistical work behind safety-weighted quality scoring, confidence calibration analysis, and ground-truth comparisons against FHIR data
- Health scoring depth: a second set of hands on body-system scoring analysis — population percentile/distribution work and outlier handling in the Databricks batching pipeline — currently handled side-of-desk by Sean alone
- RAG/embeddings evaluation (pairs with Kenan): retrieval quality analysis and embedding evaluation supporting Kenan's move toward embeddings/vector stores, without requiring him to hand off ownership
- Clinical skills library analysis: usage frequency analysis, overlap/dedup signal, and risk-tier classification accuracy for the clinical skills library
Skills
- * Real statistics and machine learning fundamentals, already in hand. Not developing basics — you can already reason about distributions, calibration, and evaluation design without a tutorial
- * Real Python ability, evidenced by your own work. A GitHub (or equivalent) that reflects code you personally wrote and can explain line-by-line — not output from a coding agent presented as personal skill
- * An understanding of systems and control flow. Coding agents are widely used now, but you can only direct one well if you understand the system you're asking it to change. You should be able to catch an agent when it's wrong, not just accept its output
- * Genuine independence in ambiguity. The problems in this role (drift thresholds, safety-weighted scoring, FHIR ground-truth comparisons) aren't fully defined yet. You should be comfortable scoping your own analysis — and just as comfortable asking a sharp, targeted question when you're stuck, and knowing who on the team to ask
- * Ambition and ownership. You treat a knowledge gap as something to close, not something to avoid, and you feel accountable for outcomes, not just tasks
- * Experience directing a coding agent as part of your own workflow (nice to have, not decisive — and its absence isn't a red flag)
- * Exposure to clinical data standards (FHIR, HL7) or health information systems
- * Coursework or project experience with embeddings, vector search, or retrieval evaluation
- * Familiarity with Databricks, Spark, or other batch-processing pipelines
- * Healthcare or other regulated-industry experience
Qualifications
Must Haves
- * Real statistics and machine learning fundamentals, already in hand. Not developing basics — you can already reason about distributions, calibration, and evaluation design without a tutorial
- * Real Python ability, evidenced by your own work. A GitHub (or equivalent) that reflects code you personally wrote and can explain line-by-line — not output from a coding agent presented as personal skill
- * An understanding of systems and control flow. Coding agents are widely used now, but you can only direct one well if you understand the system you're asking it to change. You should be able to catch an agent when it's wrong, not just accept its output
- * Genuine independence in ambiguity. The problems in this role (drift thresholds, safety-weighted scoring, FHIR ground-truth comparisons) aren't fully defined yet. You should be comfortable scoping your own analysis — and just as comfortable asking a sharp, targeted question when you're stuck, and knowing who on the team to ask
- * Ambition and ownership. You treat a knowledge gap as something to close, not something to avoid, and you feel accountable for outcomes, not just tasks
Nice to Haves
- * Experience directing a coding agent as part of your own workflow (nice to have, not decisive — and its absence isn't a red flag)
- * Exposure to clinical data standards (FHIR, HL7) or health information systems
- * Coursework or project experience with embeddings, vector search, or retrieval evaluation
- * Familiarity with Databricks, Spark, or other batch-processing pipelines
- * Healthcare or other regulated-industry experience
Benefits
- Stock options
- Incentive pay for eligible roles
- Direct senior mentorship on every workstream above