Summary
Altarum is a nonprofit organization that partners with federal and state agencies to improve health outcomes for underserved populations. The Data Engineer will design, build, deploy, and operate healthcare data pipelines, cloud platform services, and production generative AI applications, while supporting interoperability, data quality, security, and responsible AI. The role also includes technical leadership, mentoring, and knowledge sharing.
Responsibilities
- Design and operate cloud data platform components, including lakehouse storage, compute, SQL and ELT engines, and reusable transformation workflows using Databricks or comparable platforms
- Develop production Python and SQL code, versioned reusable packages, APIs, and configuration-driven frameworks with automated tests and documentation that help other teams adopt them
- Own CI/CD for data pipelines, AI applications, and supporting services, including environment configuration, automated validation, deployment, release verification, and rollback
- Implement infrastructure as code and secure deployment patterns in AWS, AWS GovCloud, Azure, Azure Government, or other approved environments
- Own ingestion, transformation, and normalization pipelines for structured, semi-structured, and unstructured healthcare and public health data, including batch and streaming workloads
- Build reusable connectors, data contracts, and curated datasets using Python, SQL, Spark or PySpark, Delta Lake, dbt, or comparable technologies
- Integrate data such as Medicaid and Medicare claims and encounters, FHIR R4, US Core, HL7 v2, social determinants of health, geospatial data, surveys, and qualitative sources as projects require
- Translate analytics, reporting, evaluation, and data science requirements into tested data models and dependable datasets with documented lineage
- Define and implement data quality rules, validation, lineage, metadata, and monitoring for owned pipelines and services
- Diagnose and resolve production failures, data quality issues, and performance bottlenecks; implement root-cause fixes and maintain operational runbooks
- Establish reliability and freshness targets with stakeholders and own monitoring, incident response, and recovery for assigned services
- Optimize compute, storage, partitioning, and pipeline execution; track cost and performance and implement guardrails
- Independently own AI application delivery from requirements and solution design through implementation, evaluation, deployment, and production support, integrating models, data, retrieval, and external services into maintainable applications
- Build retrieval-augmented generation and document-intelligence applications — including ingestion, chunking, embeddings, vector search, and retrieval — to produce source-grounded outputs from healthcare documents and other unstructured data
- Use MLflow or comparable tools to manage generative AI evaluation and tracing, evaluation datasets, custom scorers, and prompt/model versioning; assess output quality, factual accuracy, latency, and cost before and after deployment
- Develop LLM applications and agentic workflows using Python, LangChain, LlamaIndex, or comparable frameworks, with approved services such as Azure OpenAI or Azure AI Foundry
- Deploy and operate AI applications and model endpoints using Databricks Model Serving or comparable services, with governed data access, monitoring, and CI/CD
- Build reusable AI services and automation for healthcare reporting, document extraction, and knowledge retrieval, partnering with data scientists on predictive models when needed
- Apply privacy by design, encryption, least-privilege access, secrets management, and secure networking; minimize sensitive information in logs, traces, and evaluation datasets
- Implement technical controls and documentation for applicable HIPAA and 42 CFR Part 2 requirements, IRB and data-use agreement restrictions, and responsible AI practices informed by the NIST AI Risk Management Framework
- Work with data science, security, and program teams to document model behavior, evaluate fairness and explainability, and maintain model cards and other governance artifacts
- Serve as technical lead for defined projects or workstreams: break down work, coordinate contributors, review implementation choices, and help resolve blockers alongside principal engineers and project leads
- Mentor colleagues through pairing, code reviews, and technical guidance; lead workshops and create tutorials, runbooks, and examples that help teams adopt reusable data and AI tools
- Explain technical concepts and tradeoffs to varied audiences, and contribute technical materials for proposals and client engagements
Skills
- Bachelor's degree in computer science, data science, information systems, engineering, mathematics, statistics, public health informatics, or a related field, or equivalent practical experience
- Typically 3–5 years of relevant experience in data engineering, AI engineering, software engineering, analytics engineering, cloud engineering, or applied technical research, including independent delivery of production systems
- Strong Python and SQL skills, with experience building maintainable code, data models, automated tests, and reliable ingestion/transformation pipelines
- Hands-on experience with a cloud platform and a modern data platform such as Databricks or an equivalent, including distributed processing with Spark or comparable technology
- Experience using Git, CI/CD, environment-specific configuration, and automated deployment to deliver and support production workloads
- Hands-on experience building and deploying generative AI applications with LangChain, LlamaIndex, or comparable frameworks, including retrieval-augmented generation or document processing, prompt design, API integration, evaluation, and monitoring
- Ability to turn requirements into technical designs, manage defined work, troubleshoot independently, and guide collaborators through clear explanations and constructive feedback
- Candidates must be currently eligible to work in the United States; sponsorship is not available
- Databricks Certified Generative AI Engineer Associate certification strongly preferred; comparable generative AI credentials with relevant practical experience also valued
- Databricks Mosaic AI, Vector Search, Model Serving, MLflow evaluation and tracing, and Unity Catalog; experience delivering governed generative AI applications
- Healthcare and Medicaid data, quality measures, regulated reporting, FHIR or HL7, or accessible document generation
- Databricks workflows, Asset Bundles, Delta Lake, PySpark, dbt, orchestration, or infrastructure as code
- Experience leading small technical teams, mentoring colleagues or students, teaching technical subjects, or supervising applied research
Qualifications
Must Haves
- Bachelor's degree in computer science, data science, information systems, engineering, mathematics, statistics, public health informatics, or a related field, or equivalent practical experience
- Typically 3–5 years of relevant experience in data engineering, AI engineering, software engineering, analytics engineering, cloud engineering, or applied technical research, including independent delivery of production systems
- Strong Python and SQL skills, with experience building maintainable code, data models, automated tests, and reliable ingestion/transformation pipelines
- Hands-on experience with a cloud platform and a modern data platform such as Databricks or an equivalent, including distributed processing with Spark or comparable technology
- Experience using Git, CI/CD, environment-specific configuration, and automated deployment to deliver and support production workloads
- Hands-on experience building and deploying generative AI applications with LangChain, LlamaIndex, or comparable frameworks, including retrieval-augmented generation or document processing, prompt design, API integration, evaluation, and monitoring
- Ability to turn requirements into technical designs, manage defined work, troubleshoot independently, and guide collaborators through clear explanations and constructive feedback
- Candidates must be currently eligible to work in the United States; sponsorship is not available
Nice to Haves
- Databricks Certified Generative AI Engineer Associate certification strongly preferred; comparable generative AI credentials with relevant practical experience also valued
- Databricks Mosaic AI, Vector Search, Model Serving, MLflow evaluation and tracing, and Unity Catalog; experience delivering governed generative AI applications
- Healthcare and Medicaid data, quality measures, regulated reporting, FHIR or HL7, or accessible document generation
- Databricks workflows, Asset Bundles, Delta Lake, PySpark, dbt, orchestration, or infrastructure as code
- Experience leading small technical teams, mentoring colleagues or students, teaching technical subjects, or supervising applied research
Benefits
- Remote with occasional in-person collaboration days
- If you’re near one of our offices (Silver Spring, MD; or Novi, MI), you’ll be Hybrid and will join us in person one day every other month (6 times per year) for a fun, purpose-driven Collaboration Day.
- Non-local employees will join these days virtually