Open Source data engineering demo project using dbt, DuckDB, dlt, Dagster and Metabase. Two storage modes for the delta tables are supported: local and Microsoft Fabric Onelake.
-
Updated
Jul 9, 2026 - Python
Open Source data engineering demo project using dbt, DuckDB, dlt, Dagster and Metabase. Two storage modes for the delta tables are supported: local and Microsoft Fabric Onelake.
SCD2 implementation using pyspark
A modern banking data pipeline built with Dagster and DBT!
Fortune-500-grade banking analytics platform: OLTP -> medallion lakehouse -> Kimball star schema -> semantic layer -> 9-tab executive dashboard + 5 ML models (churn, fraud, segmentation, forecasting). Production-ready, governed, fully tested.
Governed ARR, retention, and AI-adoption metrics with SCD2 modeling, role-based access, and safe SQL compilation.
PostgreSQL service-desk DB for a residential complex: SCD2 dimension, validation triggers, partitioning, EXPLAIN demos and a seeded Faker data generator
Pipeline ETL MySQL en 3 couches - staging, modele en etoile avec SCD Type 2, marts analytiques. Orchestrateur Python, 18 tests de coherence inter-couches
Convention-aware integrity auditor for SCD Type 2 dimensions. 16 SQL checks covering overlaps, gaps, current-flag drift, sentinel drift and point-in-time referential breaks, with findings ranked by the fact rows and revenue at risk and idempotent repair SQL in DuckDB and Snowflake via sqlglot. 1.28M version rows audited in 3.95s.
Read-only temporal-integrity validator for SCD Type 2 dimensions. Runs seven DuckDB window-function checks (overlaps, gaps, multiple current rows, and more) and reports each violation's blast radius in keys, fact rows, and dollars.
Modern data stack reference: dbt + BigQuery + Airflow (Cloud Composer) with medallion layering, SCD2 snapshots, exposures, freshness SLAs, and 45× cost reduction via partition + cluster + incremental tuning.
This repo contains details about travel booking project executed on Databricks, Thanks
Healthcare data platform: PostgreSQL OLTP + SCD2 data warehouse + Neo4j graph.
Batch retail data lakehouse on Databricks: Delta Live Tables (bronze → silver → gold), Unity Catalog, synthetic data generator, and an executive analytics dashboard.
Pipeline de dados (medallion architecture) do dominio de delivery: dados sinteticos, MinIO, PySpark, Airflow e modelagem dimensional (star schema + SCD2)
Production-grade parameterized ETL pipeline implementing SCD Type 2 for travel booking data using Databricks, Delta Lake, and ADLS — includes data quality checks, incremental fact table build, Z-Order optimization, and SQL reporting.
Simulated core-banking ELT pipeline
Add a description, image, and links to the scd2 topic page so that developers can more easily learn about it.
To associate your repository with the scd2 topic, visit your repo's landing page and select "manage topics."