Summary
GAP Solutions, Inc. is seeking a Mid-Level Data Management Analyst to support the CDC Office of Public Health Data, Surveillance, and Technology. The role manages complex public health and healthcare data assets across their lifecycle, including ingestion, quality assurance, pipelines, stewardship, documentation, release management, and user support.
Responsibilities
- Manage CDC Data Hub data assets throughout the data lifecycle, including ingestion, release management, version control, access readiness, maintenance, and planned archival
- Perform data quality assurance, validation, and health checks to ensure data are complete, accurate, analytically consistent, and ready for production use
- Develop and maintain data extracts, transformation logic, cleaning scripts, reusable data assets, and data preparation processes using Python, SQL, R, SAS, Spark, or other approved tools
- Collaborate with data engineers to develop, maintain, monitor, and troubleshoot data pipelines and support reliable production data delivery
- Monitor and document data lineage, transformations, dependencies, schema changes, and release history
- Support data stewardship and governance activities, including metadata, data dictionaries, access requirements, usage guidance, and applicable data management policies and procedures
- Support publication, versioning, and maintenance of data products and bundles within CDC data environments and marketplaces
- Develop and maintain SOPs and technical documentation for ingestion, quality control, versioning, dissemination, and data management processes
- Serve as a liaison among CDC program staff, analysts, engineers, data vendors, and other stakeholders to clarify requirements and resolve data issues including responding to user inquiries
- Track and respond to technical assistance and user inquiries related to datasets, metadata, data dictionaries, pipelines, and data management processes
- Provide demonstrations and knowledge-sharing on data assets, analytic platforms, tools, and workflows
- Self-motivated, striving for continuous learning and willing to jump in to learn new skills and offer client base suggested process improvements
- Work independently and collaboratively across multiple concurrent data management activities and priorities
Skills
- Bachelor's degree in data science, Computer Science, Public Health, Health Informatics, Statistics, Information Systems, or a related field
- 4+ years of relevant professional experience in data management, data analytics, data engineering support, or a related field
- 1+ year of hands-on Palantir Foundry experience
- Demonstrated experience managing large, complex datasets and data pipelines
- Hands-on experience with Python, SQL, R, SAS, or comparable data programming/query languages
- Experience performing data quality assurance, validation, data preparation, transformation, and troubleshooting
- Experience with relational databases, data pipelines, data lineage, versioning, metadata, and technical documentation
- Experience working in cloud-based data or analytic environments
- Ability to obtain and maintain the required Federal Public Trust
- Databricks experience
- MPH, Master's degree, or other advanced degree in Public Health, Epidemiology, Data Science, Health Informatics, Statistics, Computer Science, or a related discipline
- Experience supporting CDC, HHS, or another Federal health/public health organization
- Experience managing healthcare or public health datasets, including EHR, claims, laboratory, pharmacy, population, or related data
- Experience with 1CDP, Palantir Foundry, Databricks, Microsoft Azure, or comparable enterprise cloud/data platforms
- Experience with data stewardship, governance, metadata management, data marketplaces, or reusable enterprise data products
- Familiarity with healthcare data structures, coding terminologies, and public health analytic use cases
- Ability to manage multiple concurrent priorities and work effectively with technical and nontechnical stakeholders
- Strong analytical, organizational, problem-solving, and communication skills
Qualifications
Must Haves
- Bachelor's degree in data science, Computer Science, Public Health, Health Informatics, Statistics, Information Systems, or a related field
- 4+ years of relevant professional experience in data management, data analytics, data engineering support, or a related field
- 1+ year of hands-on Palantir Foundry experience
- Demonstrated experience managing large, complex datasets and data pipelines
- Hands-on experience with Python, SQL, R, SAS, or comparable data programming/query languages
- Experience performing data quality assurance, validation, data preparation, transformation, and troubleshooting
- Experience with relational databases, data pipelines, data lineage, versioning, metadata, and technical documentation
- Experience working in cloud-based data or analytic environments
- Ability to obtain and maintain the required Federal Public Trust
Nice to Haves
- Databricks experience
- MPH, Master's degree, or other advanced degree in Public Health, Epidemiology, Data Science, Health Informatics, Statistics, Computer Science, or a related discipline
- Experience supporting CDC, HHS, or another Federal health/public health organization
- Experience managing healthcare or public health datasets, including EHR, claims, laboratory, pharmacy, population, or related data
- Experience with 1CDP, Palantir Foundry, Databricks, Microsoft Azure, or comparable enterprise cloud/data platforms
- Experience with data stewardship, governance, metadata management, data marketplaces, or reusable enterprise data products
- Familiarity with healthcare data structures, coding terminologies, and public health analytic use cases
- Ability to manage multiple concurrent priorities and work effectively with technical and nontechnical stakeholders
- Strong analytical, organizational, problem-solving, and communication skills
Benefits
- Fully remote work arrangement