Magpie Health Analytics logo
Magpie Health Analytics
Posted 69 days agoVerified live 17h ago

Data Engineer

Brief overview

Remote
UndergradOr in progress
4+ yrsMinimum
SQLPythonPandasPySparkNumpyGitHub ActionsJenkinsAWS GlueAWS EMRApache HadoopApache SparkAWS StepFunctionsAWS LambdaAWS S3AWS AuroraAWS SNSAWS CloudFormation

About the company

Magpie Health Analytics logo
Magpie Health Analyticsmagpiehealthanalytics.com

Magpie Health Analytics specializes in risk adjustment, population analytics, quality performance, and consumer and physician engagement.

Job description

Summary

Magpie Health Analytics is seeking a Data Engineer to support the development of cloud-native data solutions for healthcare clients. The role involves building and optimizing data pipelines and validation workflows to enhance operational efficiency and support regulatory reporting.

Responsibilities

  • Design, develop, and maintain ETL/ELT pipelines using structured and semi-structured data from relational databases, flat files, APIs, and cloud data sources
  • Collaborate with backend and architecture teams to define data transformation flows aligned with dashboard and reporting application needs
  • Design and optimize data schemas to ensure performance, integrity, and compatibility with reporting requirements
  • Develop and maintain efficient, testable, and reusable data processing scripts using Python, SQL, and cloud-native tools
  • Collaborate with DevOps, analysts, and application developers to align pipelines with system architecture, storage strategy, and reporting needs
  • Implement data quality and validation checks and document data lineage and pipeline logic for audit and reuse
  • Troubleshoot performance issues in data jobs and support data pipeline operations across environments (DEV, VAL, PROD)
  • Contribute CI/CD workflows and automation strategies to promote rapid iteration and secure deployment of data services
  • Assist in developing or maintaining data documentation, including metadata, data dictionaries, and technical user guides
  • Stay up to date with emerging technologies, techniques, and trends to inform product development and decision-making

Skills

  • Bachelor's degree in computer science, engineering, statistics, or related field
  • 4+ years of experience in a data engineering, data pipeline, or ETL/ELT development role
  • Strong proficiency in SQL, data transformation logic, and performance tuning for large datasets
  • Proficiency with Python and libraries such as Pandas, PySpark, or Numpy
  • Experience with modern version control and CI/CD practices (e.g., GitHub Actions, Jenkins)
  • Understanding of distributed computing solutions for data processing (e.g. AWS Glue, AWS EMR, Apache Hadoop, Apache Spark)
  • Experience with data pipeline orchestration frameworks or serverless tools (e.g., AWS StepFunctions, AWS Glue Workflows, and/or AWS Lambda)
  • Experience developing solutions in AWS cloud environments (S3, Lambda, Aurora, SNS, CloudFormation/CDK)
  • Experience with Snowflake, Amazon RDS, or other cloud-native data warehouses
  • Familiarity with modern data transformation tools (dbt, Dataform, or equivalent) and best practices for modular, tested SQL transformations within an ELT architecture
  • Experience supporting healthcare data systems or CMS data environments
  • AWS certifications are a plus

Qualifications

Must Haves

  • Bachelor's degree in computer science, engineering, statistics, or related field
  • 4+ years of experience in a data engineering, data pipeline, or ETL/ELT development role
  • Strong proficiency in SQL, data transformation logic, and performance tuning for large datasets
  • Proficiency with Python and libraries such as Pandas, PySpark, or Numpy
  • Experience with modern version control and CI/CD practices (e.g., GitHub Actions, Jenkins)
  • Understanding of distributed computing solutions for data processing (e.g. AWS Glue, AWS EMR, Apache Hadoop, Apache Spark)
  • Experience with data pipeline orchestration frameworks or serverless tools (e.g., AWS StepFunctions, AWS Glue Workflows, and/or AWS Lambda)

Nice to Haves

  • Experience developing solutions in AWS cloud environments (S3, Lambda, Aurora, SNS, CloudFormation/CDK)
  • Experience with Snowflake, Amazon RDS, or other cloud-native data warehouses
  • Familiarity with modern data transformation tools (dbt, Dataform, or equivalent) and best practices for modular, tested SQL transformations within an ELT architecture
  • Experience supporting healthcare data systems or CMS data environments
  • AWS certifications are a plus

Benefits

  • Remote_type: Remote (any location)

More jobs like this