Summary
Dandelion Health is building an AI training and clinical development platform that provides high-quality, de-identified healthcare data for AI development. The NLP / LLM Data Scientist will build machine learning pipelines to abstract information from multimodal healthcare data, curate AI-ready datasets, support customer AI activities, and develop analyses that improve data products.
Responsibilities
- Develop Natural Language Processing (NLP), Large Language Model (LLM) and other ML-based pipelines to abstract relevant labels from text-based healthcare data and store them in scalable data models
- Query complex source systems in a range of health data sources (e.g., EMRs, semi-structured reports, free-text clinical provider notes) to identify key data elements and create and enrich high-quality datasets for real-world evidence analyses and training AI algorithms
- Own data extraction, wrangling, labeling and QC tasks to create analytical datasets that include abstracted clinical concepts and provide a range of solutions to support customers’ AI activities
- Stay current on the latest in applied NLP and generative AI methods and proactively leverage these technologies where applicable
- Support the design, testing, validation, analysis, and merging of multimodal data structures from a wide variety of source systems
- Develop code and documentation to deliver high-quality and HIPAA-compliant data products on time to customers
- Identify and resolve problems using your knowledge, background, and troubleshooting skills
- Ensure accuracy, data integrity, and validity of data and analysis in all work
- Provide support for technical product team to advance development of the suite of data-related product offerings
- Summarize the complexity of abstraction methods, findings and recommendations into clear explanations and presentations for internal and external audiences that have a varying range of technical and clinical experience
Skills
- Advanced degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics), or B.S. with at least 5 years of professional experience
- At least 2 years of data science and machine learning experience, including building pipelines to extract and curate unstructured and semi-structured data by applying advanced machine learning and AI techniques. Prior experience with clinical and healthcare data is a strong bonus
- Fluency in Python and SQL, including fluency with ML/NLP libraries (PyTorch, Tensorflow, HuggingFace, etc.)
- Familiarity with using modern applied LLM techniques on real-world data
- Strong technical writing, editing, and communication skills, along with a collaborative mindset
- Excellent organizational skills with an ability to embrace change and effectively manage multiple projects and consistently plan work to meet deadlines
- There is occasional travel for in-person company working days on roughly a quarterly basis
- Experience working in or with startups is a plus
- Git and version control
- Familiarity with encryption methods
- Prior experience querying EDWs or databases and creating reports or analytics for healthcare data
- Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts
- Any medical ontology experience
- Any experience working with DICOM or other imaging modalities
- Experience with AWS
- Experience with publishing work in peer-reviewed journals
Qualifications
Must Haves
- Advanced degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics), or B.S. with at least 5 years of professional experience
- At least 2 years of data science and machine learning experience, including building pipelines to extract and curate unstructured and semi-structured data by applying advanced machine learning and AI techniques. Prior experience with clinical and healthcare data is a strong bonus
- Fluency in Python and SQL, including fluency with ML/NLP libraries (PyTorch, Tensorflow, HuggingFace, etc.)
- Familiarity with using modern applied LLM techniques on real-world data
- Strong technical writing, editing, and communication skills, along with a collaborative mindset
- Excellent organizational skills with an ability to embrace change and effectively manage multiple projects and consistently plan work to meet deadlines
- There is occasional travel for in-person company working days on roughly a quarterly basis
Nice to Haves
- Experience working in or with startups is a plus
- Git and version control
- Familiarity with encryption methods
- Prior experience querying EDWs or databases and creating reports or analytics for healthcare data
- Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts
- Any medical ontology experience
- Any experience working with DICOM or other imaging modalities
- Experience with AWS
- Experience with publishing work in peer-reviewed journals
Benefits
- Offers Equity
- Offers Bonus
- Remote work and flexible hours. Availability needed for meetings, which we try to keep to a healthy minimum
- Complete wellness benefits including healthcare, dental, vision, PTO, sick days and more. Ask for details
- Professional development days to build your skills
- Collegial work environment
- Academic bent towards inquiry and problem solving but start-up speed and flexibility
- Great balance of focus time to work on projects but easy to access team members to discuss issues and work collaboratively
- Dandelion is a mission-driven company that is focused on improving patient care