Summary
Washington University in St. Louis is seeking a C-BRAIN Data Engineer to build and maintain the data infrastructure supporting its AI Biomedical Research Scientist platform. The role focuses on data ingestion, pipeline development, harmonization, cloud infrastructure, data governance, quality validation, and delivery of multi-modal neurodegeneration research data for AI tools and research teams.
Responsibilities
- Designs, builds, tests, and maintains scalable data ingestion pipelines to ingest consortium member datasets from diverse sources and formats into the C-BRAIN data infrastructure
- Develops and maintains ETL/ELT workflows using tools such as Apache Spark, dbt, Airflow, or equivalent; ensure pipelines are robust, well-documented, and auditable
- Implements automated pipeline monitoring and alerting; troubleshoot and resolves pipeline failures in a timely manner
- Works collaboratively with the CTO and data science teams to understand data requirements for AI tool development and translates those requirements into technical pipeline specifications
- Maintains version control for all pipeline code and infrastructure configurations; follows software engineering best practices including code review and documentation
- Integrates and processes multi-modal data including omics (genomics, transcriptomics, proteomics), neuroimaging (PET, MRI), longitudinal clinical records, and digital pathology — reconciling differences in data type, format, spatial resolution, and dimensionality into unified analytical frameworks
- Identifies where cross-modal integration produces genuine signal versus where it introduces noise or artifact; establishes ground truth benchmarks for downstream AI use
- Manages and optimizes the C-BRAIN data infrastructure: storage accounts, computes resources, data lakes, and access controls
- Implements and maintains data access controls and permissions aligned with DUA requirements and WashU data governance policies
- Collaborates with the CTO on cloud architecture decisions; contributes to infrastructure planning for Phase 2 scale-up including foundation model compute requirements
- Monitors infrastructure costs, resource utilization, and performance; identifies and implements optimization opportunities
- Supports the deployment of C-BRAIN AI tools on cloud-based platforms; coordinates with technical teams on infrastructure requirements
- Ensures all data handling complies with DUA terms and applicable PHI de-identification requirements; implements, documents, and maintains de-identification workflows for each incoming dataset
- Uploads curated datasets to ADDI/AD Workbench and other designated repositories (NIAGADS, GP2, or equivalent) as directed; manages access controls within the platform to ensure data is accessible only by authorized users and tools
- Develops and implements data harmonization procedures to integrate datasets from multiple sources (NACC, ADNI, consortium member contributions) into a unified, analysis-ready format
- Implements data quality validation checks at ingestion and transformation stages; documents data quality issues and coordinates resolution with data providers
- Maintains comprehensive data lineage documentation: tracks data from source to consumption, documents all transformations, and ensures reproducibility
- Collaborates with research scientists and the AD, Scientific to understand scientific data requirements and ensures data products meet research use case specifications
- Aligns incoming datasets to established biomedical data standards including AD Workbench, ADDI, NIAGADS, and GP2; builds and maintains data dictionaries and metadata records for each ingested dataset
- Provides technical input on Data Use Agreements: defines technical specifications for data format, delivery method, transfer protocols, and storage requirements in coordination with the Senior Technical Product Manager
- Confirms receipt of contributed datasets, validates format and completeness against DUA specifications, and logs acceptance in the DUA register
- Flags data quality, completeness, or format issues to the Senior Technical Product Manager and CTO for follow-up with data contributors
- Supports technical aspects of the data delivery monitoring process: tracks expected deliveries, confirms receipt, and maintains data delivery logs
- Supports beta testing of data ingestion tools and provides structured feedback to development partners; maintains clear, reproducible documentation so pipeline processes can be audited and transferred
- Maintains comprehensive technical documentation for all pipelines, infrastructure configurations, and data architecture decisions in the C-BRAIN documentation repository
- Develops and maintains a C-BRAIN data catalog: documents available datasets, data dictionaries, lineage, and access procedures
- Contributes technical content to C-BRAIN progress reports, Steering Committee materials, and grants reporting as requested by the CTO or Senior Technical Product Manager
Skills
- Bachelor's degree in Computer Science, Data Science, Bioinformatics, Engineering, or a closely related field
- Three years of hands-on data engineering experience, including design and development of data pipelines and ETL/ELT workflows in a production or research environment
- Demonstrated proficiency in Python and SQL
- Experience with cloud data platforms: Microsoft Azure preferred (AWS or GCP also acceptable)
- Familiarity with cloud storage, compute, and access control management
- Experience working with complex, multi-source datasets requiring integration, harmonization, and quality validation
- Demonstrated software engineering background: production-grade Python with version control (Git), code review practices, and automated testing
- This role requires engineering discipline and the ability to build maintainable, auditable code — scripting proficiency alone is insufficient
- Experience working with at least two of the following biomedical data modalities: omics (genomics, transcriptomics, proteomics), neuroimaging (PET, MRI), digital pathology, or longitudinal clinical/EHR data
- Demonstrated experience working with neurodegeneration or Alzheimer's disease research datasets
- Familiarity with the neurodegeneration data landscape — NACC, ADNI, and/or AD/ADRD repositories — and sufficient understanding of the biological context to communicate meaningfully with research scientists
- Biomedical informatics experience without neurodegeneration domain knowledge is insufficient for this role
- Bachelor's degree
- Relevant Experience (3 Years)
- A driver's license is not required for this position
- Ability to move to on and off-campus locations
- Experience working with biomedical, clinical, or research datasets in an academic medical center, research university, or life sciences organization
- Experience with research data repositories such as ADDI, Synapse, Terra, or NACC/ADNI data platforms
- Experience with one or more of: Apache Spark, dbt, Airflow, Azure Data Factory, or equivalent ETL/ELT frameworks
- Familiarity with data governance frameworks, data use agreements, or federated data architectures
- Experience supporting AI/ML or data science teams as a data engineering partner: understanding how data products are consumed by model training and inference pipelines
- Experience with NAIRR or other research cloud computing platforms
- Experience with data catalog tools, data lineage platforms, or metadata management
- Familiarity with de-identification standards and privacy-preserving data techniques relevant to biomedical research
- Master's or PhD in Computer Science, Data Science, Bioinformatics, Biomedical Informatics, or a related field
- Familiarity with agentic AI frameworks and how curated datasets feed retrieval-augmented generation (RAG) or LLM-based co-scientist systems (e.g., LangGraph, DSPy, or equivalent)
- Experience with NLP techniques relevant to biomedical data: named entity recognition, natural language inference, or knowledge graph construction
- Knowledge of graph data structures and graph platforms (Neo4j, Amazon Neptune, or equivalent) for representing multi-modal biomedical relationships
- Track record of cross-disciplinary collaboration between computational and experimental or clinical teams
- Metadata Repository
- Cloud Computing Platform
- Computer Science
- Data Engineering
- Federated Identity Management
- Research Databases
- Master's degree, PhD or terminal degree or combination of education and experience may substitute for minimum education
- Academic Disciplines, AI Frameworks, Apache Airflow, Apache Spark, Apache Synapse, Azure Data Factory, Bioinformatics, Biomedical Data, Biomedical Informatics, Catalog Management, Data ETL, Data Governance Framework, Data Lineage, Data Management, Data Management Platforms, Data Pipelines, Data Privacy Protection, Data Science, Data Security Management, Data Standards, dbt Core, Generative AI, Graph Databases, Machine Learning (ML), Medical Centers
Qualifications
Must Haves
- Bachelor's degree in Computer Science, Data Science, Bioinformatics, Engineering, or a closely related field
- Three years of hands-on data engineering experience, including design and development of data pipelines and ETL/ELT workflows in a production or research environment
- Demonstrated proficiency in Python and SQL
- Experience with cloud data platforms: Microsoft Azure preferred (AWS or GCP also acceptable)
- Familiarity with cloud storage, compute, and access control management
- Experience working with complex, multi-source datasets requiring integration, harmonization, and quality validation
- Demonstrated software engineering background: production-grade Python with version control (Git), code review practices, and automated testing
- This role requires engineering discipline and the ability to build maintainable, auditable code — scripting proficiency alone is insufficient
- Experience working with at least two of the following biomedical data modalities: omics (genomics, transcriptomics, proteomics), neuroimaging (PET, MRI), digital pathology, or longitudinal clinical/EHR data
- Demonstrated experience working with neurodegeneration or Alzheimer's disease research datasets
- Familiarity with the neurodegeneration data landscape — NACC, ADNI, and/or AD/ADRD repositories — and sufficient understanding of the biological context to communicate meaningfully with research scientists
- Biomedical informatics experience without neurodegeneration domain knowledge is insufficient for this role
- Bachelor's degree
- Relevant Experience (3 Years)
- A driver's license is not required for this position
- Ability to move to on and off-campus locations
Nice to Haves
- Experience working with biomedical, clinical, or research datasets in an academic medical center, research university, or life sciences organization
- Experience with research data repositories such as ADDI, Synapse, Terra, or NACC/ADNI data platforms
- Experience with one or more of: Apache Spark, dbt, Airflow, Azure Data Factory, or equivalent ETL/ELT frameworks
- Familiarity with data governance frameworks, data use agreements, or federated data architectures
- Experience supporting AI/ML or data science teams as a data engineering partner: understanding how data products are consumed by model training and inference pipelines
- Experience with NAIRR or other research cloud computing platforms
- Experience with data catalog tools, data lineage platforms, or metadata management
- Familiarity with de-identification standards and privacy-preserving data techniques relevant to biomedical research
- Master's or PhD in Computer Science, Data Science, Bioinformatics, Biomedical Informatics, or a related field
- Familiarity with agentic AI frameworks and how curated datasets feed retrieval-augmented generation (RAG) or LLM-based co-scientist systems (e.g., LangGraph, DSPy, or equivalent)
- Experience with NLP techniques relevant to biomedical data: named entity recognition, natural language inference, or knowledge graph construction
- Knowledge of graph data structures and graph platforms (Neo4j, Amazon Neptune, or equivalent) for representing multi-modal biomedical relationships
- Track record of cross-disciplinary collaboration between computational and experimental or clinical teams
- Metadata Repository
- Cloud Computing Platform
- Computer Science
- Data Engineering
- Federated Identity Management
- Research Databases
- Master's degree, PhD or terminal degree or combination of education and experience may substitute for minimum education
- Academic Disciplines, AI Frameworks, Apache Airflow, Apache Spark, Apache Synapse, Azure Data Factory, Bioinformatics, Biomedical Data, Biomedical Informatics, Catalog Management, Data ETL, Data Governance Framework, Data Lineage, Data Management, Data Management Platforms, Data Pipelines, Data Privacy Protection, Data Science, Data Security Management, Data Standards, dbt Core, Generative AI, Graph Databases, Machine Learning (ML), Medical Centers
Benefits
- Up to 22 days of vacation, 10 recognized holidays, and sick time.
- Competitive health insurance packages with priority appointments and lower copays/coinsurance.
- Free Metro transit U-Pass for eligible employees.
- Eligible employees receive a defined contribution (403(b)) Retirement Savings Plan, which combines employee contributions and university contributions starting at 7%.
- Wellness challenges, annual health screenings, mental health resources, mindfulness programs and courses, employee assistance program (EAP), financial resources, access to dietitians, and more.
- 4 weeks of caregiver leave to bond with your new child.
- Family care resources are available for continued childcare needs and adult care.
- WashU covers the cost of tuition for you and your family, including dependent undergraduate-level college tuition up to 100% at WashU and 40% elsewhere after seven years with us.