Summary
MetaPhase is a mission-focused consultancy that values authenticity, collaboration, and practical innovation for its clients. The Data Engineer – Databricks will design, develop, test, deploy, operate, and continuously improve data pipelines and data products in Databricks, while supporting governance, troubleshooting, platform operations, and enterprise data delivery.
Responsibilities
- Develop, test, deploy, and maintain batch and streaming data pipelines using Databricks, Python, SQL, Apache Spark, and Delta Lake
- Build and enhance ingestion, transformation, validation, and publishing processes that move data from source systems into governed data products and analytics-ready datasets
- Implement approved data models, data-quality rules, metadata, and documentation in accordance with established architecture and governance standards
- Configure and maintain Databricks notebooks, workflows, jobs, compute resources, and related deployment artifacts
- Participate in code reviews, peer testing, release preparation, defect remediation, and CI/CD activities
- Monitor pipeline performance, job execution, data-quality results, and platform alerts; troubleshoot issues and support resolution of production incidents
- Collaborate with architects, analysts, data owners, and other engineers to clarify requirements, identify dependencies, and deliver iterative improvements
- Maintain technical documentation for pipelines, data sources, transformations, interfaces, test results, and operating procedures
Skills
- Bachelor's degree in a technical discipline and three or more years of relevant experience in data engineering, software engineering, analytics engineering, or a related field
- Demonstrated proficiency in Python and SQL, with experience developing, debugging, and maintaining ETL/ELT processes and data-processing code
- One or more years of hands-on experience with Databricks, Apache Spark, or a comparable cloud data-engineering platform
- Experience working with structured and/or unstructured data sources, data validation, source-to-target mapping, and production-support activities
- Ability to obtain a U.S. Public Trust suitability determination
- U.S. Citizenship Required
- Databricks Certified Data Engineer Associate certification
- Experience with Databricks capabilities such as Delta Lake, Auto Loader, Databricks SQL, Lakeflow Jobs, Unity Catalog, or streaming data pipelines
- Familiarity with Git-based version control, code reviews, automated testing, CI/CD, and Agile delivery practices
- Experience supporting data governance activities, including metadata documentation, data-quality checks, lineage, and access-control implementation
- Experience with AWS, Azure, or Google Cloud data services and cloud-based integrations
- Experience supporting regulated, public-sector, or security-sensitive data environments
- Additional Databricks certifications (e.g., Databricks Machine Learning Engineer Associate or Professional, Databricks Generative AI Engineer Associate)
Qualifications
Must Haves
- Bachelor's degree in a technical discipline and three or more years of relevant experience in data engineering, software engineering, analytics engineering, or a related field
- Demonstrated proficiency in Python and SQL, with experience developing, debugging, and maintaining ETL/ELT processes and data-processing code
- One or more years of hands-on experience with Databricks, Apache Spark, or a comparable cloud data-engineering platform
- Experience working with structured and/or unstructured data sources, data validation, source-to-target mapping, and production-support activities
- Ability to obtain a U.S. Public Trust suitability determination
- U.S. Citizenship Required
Nice to Haves
- Databricks Certified Data Engineer Associate certification
- Experience with Databricks capabilities such as Delta Lake, Auto Loader, Databricks SQL, Lakeflow Jobs, Unity Catalog, or streaming data pipelines
- Familiarity with Git-based version control, code reviews, automated testing, CI/CD, and Agile delivery practices
- Experience supporting data governance activities, including metadata documentation, data-quality checks, lineage, and access-control implementation
- Experience with AWS, Azure, or Google Cloud data services and cloud-based integrations
- Experience supporting regulated, public-sector, or security-sensitive data environments
- Additional Databricks certifications (e.g., Databricks Machine Learning Engineer Associate or Professional, Databricks Generative AI Engineer Associate)
Benefits
- Generous PTO
- Federal holidays
- Parental leave
- Comprehensive health coverage (medical, dental, vision, life, and disability)
- 401(k) with company match
- FSA/HSA options
- Commuter benefits
- Remote work