Summary
Dandelion is a health tech company focused on building the world’s largest AI training and clinical development platform. They are seeking an experienced Analytics Engineer to build, maintain, and optimize the transformation layer of their ELT pipeline, ensuring the delivery of clean, analysis-ready datasets for clients and internal teams.
Responsibilities
- Querying complex source systems in a range of health data sources (e.g., EMRs, ECG data, DICOM data) to identify and map data elements in order to create high-quality datasets to support training AI algorithms and conducting in-depth analyses of patient journeys
- Developing and maintaining data transformation pipelines using dbt and Snowflake, ensuring data quality, lineage, and transparency
- Harmonizing multimodal data from multiple health systems into a unified ontological layer ready for use by data scientists and AI developers, with input from clinical experts across the company
- Performing complex data extraction, manipulation, and summarization of large databases to create analytical data models
- Implementing data quality tests, CI/CD and version control best practices for analytics codebases throughout the data modeling pipeline
- Supporting software engineers in optimizing processes and technical solutions for de-identification and ETL of clinical data from disparate health system cloud environments to Dandelion Health’s data warehouse
- Working with experts in natural language processing (NLP) to structure and model discrete clinical concepts data that have been abstracted from unstructured and semi-structured text data
- Collaborating with technical and clinical subject matter experts and Dandelion customers to translate research and model requirements into engineering solutions
Skills
- Bachelor's degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics)
- 3+ years of experience in analytics or data engineering roles where you are responsible for hands-on cleaning and structuring clinical and electronic health record data
- Strong dbt and SQL proficiency: writing reusable Jinja macros, implementing custom tests, and comfortable with multi-environment deployments (dev/staging/prod)
- 1+ years of experience with extracting, curating, and analyzing data created within the HIT and healthcare delivery ecosystem (e.g., EMR, claims, registry); this may include knowledge of the roles of data exchange and content standards (e.g., FHIR, CDA, CQL) and clinical terminology standards (e.g., ICD, CPT, LOINC, SNOMED-CT, NDC, RxNorm)
- Comfort working in a cloud environment (ideally Snowflake and/or AWS)
- Proficiency with Git and version control workflows
- Is an advocate for bringing software engineering best practices (modularity, unit testing, concise and well-documented code) to a data science team
- Comfortable with ambiguity and creative problem-solving in a fast-paced, rapidly growing startup environment
- Excellent communication skills to advocate for the value of robust analytics engineering solutions and translate customer needs into data solutions
- Passion for improving healthcare and building infrastructure that makes research and clinical AI safer, faster, and more reliable
- Practical experience integrating LLM tools (e.g., Snowflake Cortex, Amazon Bedrock, Azure OpenAI Service) and frameworks (RAG, agent workflows) into data pipelines and product
- Proficiency in Python for data cleaning and model feature engineering
- Experience working with AI/ML modeling teams and a conceptual grasp of cloud model deployments and evaluation
- Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts
- Any medical ontology experience
- Any experience working with DICOM or other imaging modalities
- Experience with OMOP common data model
- Comfortable with agile tools such as Jira, Linear
- Familiarity using dbt cloud
Qualifications
Must Haves
- Bachelor's degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics)
- 3+ years of experience in analytics or data engineering roles where you are responsible for hands-on cleaning and structuring clinical and electronic health record data
- Strong dbt and SQL proficiency: writing reusable Jinja macros, implementing custom tests, and comfortable with multi-environment deployments (dev/staging/prod)
- 1+ years of experience with extracting, curating, and analyzing data created within the HIT and healthcare delivery ecosystem (e.g., EMR, claims, registry); this may include knowledge of the roles of data exchange and content standards (e.g., FHIR, CDA, CQL) and clinical terminology standards (e.g., ICD, CPT, LOINC, SNOMED-CT, NDC, RxNorm)
- Comfort working in a cloud environment (ideally Snowflake and/or AWS)
- Proficiency with Git and version control workflows
- Is an advocate for bringing software engineering best practices (modularity, unit testing, concise and well-documented code) to a data science team
- Comfortable with ambiguity and creative problem-solving in a fast-paced, rapidly growing startup environment
- Excellent communication skills to advocate for the value of robust analytics engineering solutions and translate customer needs into data solutions
- Passion for improving healthcare and building infrastructure that makes research and clinical AI safer, faster, and more reliable
Nice to Haves
- Practical experience integrating LLM tools (e.g., Snowflake Cortex, Amazon Bedrock, Azure OpenAI Service) and frameworks (RAG, agent workflows) into data pipelines and product
- Proficiency in Python for data cleaning and model feature engineering
- Experience working with AI/ML modeling teams and a conceptual grasp of cloud model deployments and evaluation
- Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts
- Any medical ontology experience
- Any experience working with DICOM or other imaging modalities
- Experience with OMOP common data model
- Comfortable with agile tools such as Jira, Linear
- Familiarity using dbt cloud
Benefits
- Remote work and flexible hours. Availability needed for meetings, which we try to keep to a healthy minimum
- Complete wellness benefits including healthcare, dental, vision, PTO, sick days and more. Ask for details
- Professional development days to build your skills
- Collegial work environment
- Academic bent towards inquiry and problem solving but start-up speed and flexibility
- Great balance of focus time to work on projects but easy to access team members to discuss issues and work collaboratively
- Dandelion is a mission-driven company that is focused on improving patient care