Summary
Beacon Talent is recruiting for a high-growth, mission-driven healthcare AI and clinical data platform company that builds large-scale, de-identified datasets for responsible AI development. The Analytics Engineer will own the data modeling lifecycle, including transformation pipelines, data quality, testing, documentation, and performance optimization. The role will also integrate multimodal clinical data and collaborate with technical and clinical experts to support AI training and patient journey analysis.
Responsibilities
- Query complex source systems across EMR, ECG, and DICOM data to identify and map data elements supporting AI training and patient journey analyses
- Develop and maintain data transformation pipelines using dbt and Snowflake, ensuring data quality, lineage, and transparency
- Harmonize multimodal data from multiple health systems into a unified ontological layer, with input from clinical experts
- Perform complex data extraction, manipulation, and summarization to create analytical data models
- Implement data quality tests, CI/CD, and version control best practices across the data modeling pipeline
- Support software engineers in optimizing de-identification and ETL processes across disparate health system cloud environments
- Work with NLP experts to structure and model discrete clinical concepts abstracted from unstructured text
- Collaborate with technical and clinical SMEs and customers to translate research and model requirements into engineering solutions
Skills
- Bachelor's degree in a quantitative field (Data Science, Biomedical Informatics, Computer Science, Biostatistics)
- 3+ years in analytics or data engineering roles with hands-on cleaning and structuring of clinical/EHR data
- Strong dbt and SQL proficiency: reusable Jinja macros, custom tests, multi-environment deployments
- 1+ years extracting, curating, and analyzing HIT/healthcare delivery data (EMR, claims, registry); familiarity with FHIR, CDA, CQL, and clinical terminology standards (ICD, CPT, LOINC, SNOMED-CT, NDC, RxNorm) a plus
- Proficiency with Git and version control workflows
- Advocate for software engineering best practices (modularity, unit testing, clean documentation) within a data science team
- Comfortable with ambiguity in a fast-paced, early-stage startup environment
- Excellent communication skills, able to translate customer needs into data solutions
- Comfort in a cloud environment (Snowflake and/or AWS preferred)
Qualifications
Must Haves
- Bachelor's degree in a quantitative field (Data Science, Biomedical Informatics, Computer Science, Biostatistics)
- 3+ years in analytics or data engineering roles with hands-on cleaning and structuring of clinical/EHR data
- Strong dbt and SQL proficiency: reusable Jinja macros, custom tests, multi-environment deployments
- 1+ years extracting, curating, and analyzing HIT/healthcare delivery data (EMR, claims, registry); familiarity with FHIR, CDA, CQL, and clinical terminology standards (ICD, CPT, LOINC, SNOMED-CT, NDC, RxNorm) a plus
- Proficiency with Git and version control workflows
- Advocate for software engineering best practices (modularity, unit testing, clean documentation) within a data science team
- Comfortable with ambiguity in a fast-paced, early-stage startup environment
- Excellent communication skills, able to translate customer needs into data solutions
Nice to Haves
- Comfort in a cloud environment (Snowflake and/or AWS preferred)
Benefits
- Remote work and flexible hours, with meetings kept to a healthy minimum
- Comprehensive wellness benefits: healthcare, dental, vision, PTO, sick days
- Professional development days
- Collegial, academically-minded culture with startup speed and flexibility
- Strong balance of focus time and easy access to collaborators
- Mission-driven company focused on improving patient care through AI