Summary
Vālenz Health® advances healthcare through technology and clinical expertise to improve cost, quality, access, and utilization. The Data Engineer I will develop and maintain scalable cloud-based data pipelines, transform and integrate healthcare data, and support the evolution of the lakehouse architecture while partnering with analytics and cross-functional teams.
Responsibilities
- Develop, maintain, and enhance data pipelines that ingest, transform, validate, and integrate data from multiple internal and external sources
- Support the migration of on-premise SQL Server data systems to a cloud-based lakehouse architecture using Azure Databricks and Delta Lake
- Develop and maintain ETL/ELT processes using SQL, Python, PySpark, and Spark SQL
- Apply established lakehouse and Delta Lake architecture standards, including schema enforcement, data quality controls, ACID transactions, and data versioning
- Build and maintain data models that support analytics, reporting, and data warehousing needs
- Implement data quality checks, validation processes, and monitoring to ensure the accuracy, completeness, and reliability of data pipelines
- Orchestrate and monitor data workflows using Databricks Workflows, Apache Airflow, or similar technologies
- Troubleshoot pipeline failures, data quality issues, and performance concerns, identifying root causes and implementing appropriate solutions
- Support the optimization of data pipelines for performance, scalability, reliability, and cost efficiency
- Collaborate on CI/CD practices for data engineering solutions, including source control, testing, deployment, and version management
- Partner with data analysts, data scientists, and business stakeholders to understand data requirements and develop supporting pipelines and data structures
- Document data pipelines, transformations, technical processes, and data structures to support maintainability and knowledge sharing
- Participate in code reviews, testing, and Agile development processes to support consistent engineering standards and continuous improvement
- Stay current on data engineering technologies and contribute ideas for improving the organization's evolving data platform
- Perform other duties as assigned
Skills
- Bachelor's degree in Computer Science, Engineering, Data Science, Mathematics, Statistics, or a related quantitative or technical field, or equivalent practical experience
- 1+ years of experience in data engineering, software development, data analytics, or a related role involving similar technical responsibilities
- Working knowledge of SQL and Python, with experience using these technologies to manipulate, transform, or process data
- Experience developing or supporting ETL/ELT processes, data pipelines, data integrations, or similar data engineering solutions
- Understanding of relational databases, data modeling concepts, and data warehousing principles
- Ability to troubleshoot data and technical issues, investigate root causes, and develop effective solutions
- Strong attention to detail with a commitment to data accuracy, quality, and validation
- Ability to work with complex or incomplete data and navigate ambiguity to reach reliable conclusions
- Strong organizational and time management skills with the ability to manage multiple priorities and assignments
- Effective written and verbal communication skills, including the ability to explain technical concepts and requirements to technical and non-technical stakeholders
- You'll need a quiet workspace that is free from distractions
- Reliable internet connection—if you can use streaming services, you're good to go!
- Adherence to company security protocols, including the use of VPNs, secure passwords, and company-approved devices/software
- You must be US based, in a location where you can work effectively and comply with company policies such as HIPAA
- Experience with Azure Databricks, Spark, PySpark, Spark SQL, or Delta Lake
- Experience with Microsoft Azure services such as Azure Data Lake Storage, Blob Storage, or Synapse Analytics
- Experience with pipeline orchestration tools such as Databricks Workflows, Apache Airflow, or similar technologies
- Familiarity with Git, Azure DevOps, CI/CD practices, automated testing, and source control
- Experience working with healthcare data, including medical claims, eligibility, provider network, pharmacy claims, or similar datasets
- Exposure to cloud data migrations or modernization of traditional relational database and ETL environments
- Familiarity with batch and streaming data processing concepts
- Understanding of data quality frameworks, monitoring, and data governance practices
Qualifications
Must Haves
- Bachelor's degree in Computer Science, Engineering, Data Science, Mathematics, Statistics, or a related quantitative or technical field, or equivalent practical experience
- 1+ years of experience in data engineering, software development, data analytics, or a related role involving similar technical responsibilities
- Working knowledge of SQL and Python, with experience using these technologies to manipulate, transform, or process data
- Experience developing or supporting ETL/ELT processes, data pipelines, data integrations, or similar data engineering solutions
- Understanding of relational databases, data modeling concepts, and data warehousing principles
- Ability to troubleshoot data and technical issues, investigate root causes, and develop effective solutions
- Strong attention to detail with a commitment to data accuracy, quality, and validation
- Ability to work with complex or incomplete data and navigate ambiguity to reach reliable conclusions
- Strong organizational and time management skills with the ability to manage multiple priorities and assignments
- Effective written and verbal communication skills, including the ability to explain technical concepts and requirements to technical and non-technical stakeholders
- You'll need a quiet workspace that is free from distractions
- Reliable internet connection—if you can use streaming services, you're good to go!
- Adherence to company security protocols, including the use of VPNs, secure passwords, and company-approved devices/software
- You must be US based, in a location where you can work effectively and comply with company policies such as HIPAA
Nice to Haves
- Experience with Azure Databricks, Spark, PySpark, Spark SQL, or Delta Lake
- Experience with Microsoft Azure services such as Azure Data Lake Storage, Blob Storage, or Synapse Analytics
- Experience with pipeline orchestration tools such as Databricks Workflows, Apache Airflow, or similar technologies
- Familiarity with Git, Azure DevOps, CI/CD practices, automated testing, and source control
- Experience working with healthcare data, including medical claims, eligibility, provider network, pharmacy claims, or similar datasets
- Exposure to cloud data migrations or modernization of traditional relational database and ETL environments
- Familiarity with batch and streaming data processing concepts
- Understanding of data quality frameworks, monitoring, and data governance practices
Benefits
- Fully remote position
- All necessary equipment provided
- Generously subsidized company-sponsored Medical, Dental, and Vision insurance, with access to services through our own products, Healthcare Blue Book and KISx Card.
- Spending account options: HSA, FSA, and DCFSA
- 401K with company match and immediate vesting
- Flexible working environment
- Generous Paid Time Off to include vacation, sick leave, and paid holidays
- Employee Assistance Program that includes professional counseling, referrals, and additional services
- Paid maternity and paternity leave
- Pet insurance
- Employee discounts on phone plans, car rentals and computers
- Community giveback opportunities, including paid time off for philanthropic endeavors