Summary
VWO is seeking a Big Data Engineer to design, build, and operate large-scale data pipelines and analytical infrastructure for analytics, reporting, and data-driven products. The role owns data workflows from ingestion through warehousing, orchestration, quality, and observability, with primary responsibility for ClickHouse and collaboration across data science, analytics, product, and engineering teams.
Responsibilities
- Design, build, and maintain robust **batch and streaming data pipelines** that ingest data from multiple sources into analytical data stores
- Build and operate **Apache Airflow DAGs**, including scheduling, dependencies, retries, backfills, idempotency, concurrency, and failure handling
- Develop analytics-ready datasets using **dbt**, following well-structured staging, intermediate, and mart layers with appropriate tests, documentation, and incremental models
- Own **ClickHouse** as the primary analytical store, including:
+ Schema and table design using the MergeTree family of engines
+ Partitioning and sorting/primary key strategies
+ Materialized views
+ Distributed and replicated table architectures
+ Query and memory optimization
+ High-volume data ingestion and performance tuning
- Work with **BigQuery** where cloud data-warehouse patterns are appropriate, including data modeling and query/cost optimization
- Design and operate **NoSQL and key-value data stores**, including Bigtable, DynamoDB, and Redis, based on specific access patterns and performance requirements
- Build and maintain **data-quality frameworks** covering validation, testing, freshness, completeness, reconciliation, and anomaly detection
- Implement **monitoring, alerting, structured logging, and observability** for data pipelines and services
- Own pipeline SLAs, incident response, troubleshooting, and root-cause analysis
- Manage **backfills, safe re-runs, schema evolution, and data migrations** while minimizing downstream impact
- Build reproducible, containerized environments using **Docker** and contribute to **CI/CD and Infrastructure as Code** practices
- Partner with analysts, data scientists, product managers, and product engineers to translate business and technical requirements into scalable data models and pipelines
- Continuously improve pipeline reliability, scalability, performance, and infrastructure cost efficiency
Skills
- * **4-6 years of experience** in data engineering or a closely related field, with strong hands-on production experience
- * Expert-level **SQL** and strong **Python** skills, with experience writing production-grade, maintainable, and well-tested code
- * Strong hands-on experience with **Apache Airflow** or a comparable workflow orchestration platform, including DAG design, scheduling, retries, backfills, dependency management, idempotency, and concurrency
- * Hands-on experience with **dbt** or a comparable transformation/ELT framework, including modular models, testing, source management, documentation, and incremental processing
- * **Expert-level production experience with ClickHouse**. This is a core requirement and should include:
- + MergeTree engine family
- + Partitioning and primary/sorting keys
- + Materialized views
- + Distributed and replicated tables
- + Query and memory optimization
- + High-volume ingestion and performance tuning
- * Production experience with a **cloud-based columnar/OLAP warehouse**, such as BigQuery, including data modeling and performance/cost optimization
- * Hands-on experience with **NoSQL and key-value databases**, such as Bigtable, DynamoDB, and Redis, including data modeling, partition/key design, access patterns, and caching strategies
- * Strong understanding of **ETL/ELT and dimensional/layered data modeling** principles
- * Experience with **GCP and/or AWS** and familiarity with **Docker, Git, and CI/CD**
- * Strong focus on **data quality, reliability, observability, and correctness**
- * Strong ownership, problem-solving, and communication skills, with the ability to collaborate effectively across engineering, analytics, data science, and product teams
- * Experience with **search platforms** such as Elasticsearch, OpenSearch, Apache Solr, or Vespa, including indexing pipelines, schema design, and relevance/performance tuning
- * Experience with **Aerospike** or other high-performance, low-latency distributed key-value/NoSQL systems
- * Experience building **streaming and event-driven pipelines** using Kafka, Pub/Sub, or similar technologies
- * Experience with **Change Data Capture (CDC)** patterns and technologies
- * Experience with **Apache Spark** and data-lake architectures using object storage such as GCS or S3
- * Experience with **Terraform** or other Infrastructure as Code tools and **Kubernetes**
- * Experience with data-quality and observability tools such as **Great Expectations, Soda, Monte Carlo, or advanced dbt testing**
- * Understanding of **data platform cost optimization / FinOps** practices
- * Experience handling **high-volume e-commerce, product catalog, behavioral, or event data**
- * Familiarity with a second backend programming language, particularly **Go**
Qualifications
Must Haves
- * **4-6 years of experience** in data engineering or a closely related field, with strong hands-on production experience
- * Expert-level **SQL** and strong **Python** skills, with experience writing production-grade, maintainable, and well-tested code
- * Strong hands-on experience with **Apache Airflow** or a comparable workflow orchestration platform, including DAG design, scheduling, retries, backfills, dependency management, idempotency, and concurrency
- * Hands-on experience with **dbt** or a comparable transformation/ELT framework, including modular models, testing, source management, documentation, and incremental processing
- * **Expert-level production experience with ClickHouse**. This is a core requirement and should include:
- + MergeTree engine family
- + Partitioning and primary/sorting keys
- + Materialized views
- + Distributed and replicated tables
- + Query and memory optimization
- + High-volume ingestion and performance tuning
- * Production experience with a **cloud-based columnar/OLAP warehouse**, such as BigQuery, including data modeling and performance/cost optimization
- * Hands-on experience with **NoSQL and key-value databases**, such as Bigtable, DynamoDB, and Redis, including data modeling, partition/key design, access patterns, and caching strategies
- * Strong understanding of **ETL/ELT and dimensional/layered data modeling** principles
- * Experience with **GCP and/or AWS** and familiarity with **Docker, Git, and CI/CD**
- * Strong focus on **data quality, reliability, observability, and correctness**
- * Strong ownership, problem-solving, and communication skills, with the ability to collaborate effectively across engineering, analytics, data science, and product teams
Nice to Haves
- * Experience with **search platforms** such as Elasticsearch, OpenSearch, Apache Solr, or Vespa, including indexing pipelines, schema design, and relevance/performance tuning
- * Experience with **Aerospike** or other high-performance, low-latency distributed key-value/NoSQL systems
- * Experience building **streaming and event-driven pipelines** using Kafka, Pub/Sub, or similar technologies
- * Experience with **Change Data Capture (CDC)** patterns and technologies
- * Experience with **Apache Spark** and data-lake architectures using object storage such as GCS or S3
- * Experience with **Terraform** or other Infrastructure as Code tools and **Kubernetes**
- * Experience with data-quality and observability tools such as **Great Expectations, Soda, Monte Carlo, or advanced dbt testing**
- * Understanding of **data platform cost optimization / FinOps** practices
- * Experience handling **high-volume e-commerce, product catalog, behavioral, or event data**
- * Familiarity with a second backend programming language, particularly **Go**