Summary
CoreView provides a cloud-native platform for Microsoft 365 security, governance, administration, and tenant resilience. The Data Engineer will establish and operate reliable data pipelines and governance practices, own customer usage and adoption analytics, and consolidate data sources using Databricks, SQL, and Python.
Responsibilities
- Take operational ownership of the usage and adoption pipelines and their alerting
- Reconcile figures derived from our product-usage tool
- Separate development and production environments and mature CI/CD so deployment is repeatable and rollback is reliable
- Build and maintain Bronze-to-Gold pipelines in Databricks, Unity Catalog metric views, and the permissions/row-level-security model behind them
- Extend automated data quality checks across the underlying gold-layer tables and validate outputs against existing Power BI reports
- Co-own the scoping of the next set of usage and adoption signals with Product and Post-Sales: prioritising upsell signals and the ROI narrative and confirming what needs new instrumentation
- Handle the realities of business system data: schema drift, inconsistent field naming, soft deletes, incremental loads
- Lead and be responsible for the consolidation of additional data sources — including support tickets, audit logs, CRM, and other customer-facing systems — while defining the feasibility, effort, cost, and implementation approach required to support a future unified data platform
- Establish documented pipelines, an access model for raw and transformed data, and a working intake process for cross-department data requests
- Act as a credible point of contact for CoreView's Databricks estate
- Bring automation and AI-assisted tooling into the day-to-day work — for example automated monitoring and anomaly detection on metric movements — as the platform matures
Skills
- 3 to 5 years' professional experience building and operating production data pipelines (ETL/ELT) in a commercial setting
- Hands-on Databricks proficiency: Unity Catalog, Spark Declarative Pipelines and metric views, including the permissions / row-level-security model behind them
- Strong SQL and Python applied to data engineering work — building, testing and supporting pipeline code through to production
- Experience with data transformation/orchestration tooling such as Spark Declarative Pipelines, dbt, or an equivalent
- Solid grounding in data modelling (dimensional/star-schema approaches) for analytics consumption
- Git-based version control and CI/CD practices applied to data pipelines, not just application code
- Experience ingesting data from a range of source systems — CRM (Salesforce, HubSpot), product usage and telemetry (CoreView's own platform, UserPilot, or equivalents such as Pendo, Amplitude or Mixpanel), and other business systems
- Comfortable operating with limited hand-holding: implementing, documenting and running to an agreed direction rather than waiting for detailed instruction
- Clear written and verbal communication: you'll be the named contact for commercial teams asking questions about the numbers
- Power BI or an equivalent BI tool, including validating a migrated report against a legacy one
- Prior experience implementing data governance, a metric catalogue, or an access model from a low base of maturity
- Background in B2B SaaS
- Experience using AI coding assistants (GitHub Copilot, Cursor, Claude) as part of a production engineering workflow
Qualifications
Must Haves
- 3 to 5 years' professional experience building and operating production data pipelines (ETL/ELT) in a commercial setting
- Hands-on Databricks proficiency: Unity Catalog, Spark Declarative Pipelines and metric views, including the permissions / row-level-security model behind them
- Strong SQL and Python applied to data engineering work — building, testing and supporting pipeline code through to production
- Experience with data transformation/orchestration tooling such as Spark Declarative Pipelines, dbt, or an equivalent
- Solid grounding in data modelling (dimensional/star-schema approaches) for analytics consumption
- Git-based version control and CI/CD practices applied to data pipelines, not just application code
- Experience ingesting data from a range of source systems — CRM (Salesforce, HubSpot), product usage and telemetry (CoreView's own platform, UserPilot, or equivalents such as Pendo, Amplitude or Mixpanel), and other business systems
- Comfortable operating with limited hand-holding: implementing, documenting and running to an agreed direction rather than waiting for detailed instruction
- Clear written and verbal communication: you'll be the named contact for commercial teams asking questions about the numbers
Nice to Haves
- Power BI or an equivalent BI tool, including validating a migrated report against a legacy one
- Prior experience implementing data governance, a metric catalogue, or an access model from a low base of maturity
- Background in B2B SaaS
- Experience using AI coding assistants (GitHub Copilot, Cursor, Claude) as part of a production engineering workflow
Benefits
- Remote work in the United States, with a hybrid option in Washington, DC