Open Source data engineering demo project using dbt, DuckDB, dlt, Dagster and Metabase. Two storage modes for the delta tables are supported: local and Microsoft Fabric Onelake.
-
Updated
Jul 9, 2026 - Python
Open Source data engineering demo project using dbt, DuckDB, dlt, Dagster and Metabase. Two storage modes for the delta tables are supported: local and Microsoft Fabric Onelake.
SCD2 implementation using pyspark
A modern banking data pipeline built with Dagster and DBT!
Fortune-500-grade banking analytics platform: OLTP -> medallion lakehouse -> Kimball star schema -> semantic layer -> 9-tab executive dashboard + 5 ML models (churn, fraud, segmentation, forecasting). Production-ready, governed, fully tested.
Governed ARR, retention, and AI-adoption metrics with SCD2 modeling, role-based access, and safe SQL compilation.
Legacy HR and payroll data migrated into a Workday-shaped target: staged, cleansed, validated, loaded and reconciled to the cent, with every rejection and change traced to a numbered rule. Pure Python + SQLite, zero dependencies, with an interactive replay of the run.
PostgreSQL service-desk DB for a residential complex: SCD2 dimension, validation triggers, partitioning, EXPLAIN demos and a seeded Faker data generator
This is a data engineering pipeline built on Databricks + Delta Lake + PySpark that ingests travel booking and customer master data, applies SCD Type 2 logic, and delivers analytics-ready tables. It includes data quality enforcement, dimension versioning, fact aggregation, and performance tuning.
Pipeline ETL MySQL en 3 couches - staging, modele en etoile avec SCD Type 2, marts analytiques. Orchestrateur Python, 18 tests de coherence inter-couches
Analytics data warehouse for a production livestock ERP — dbt + BigQuery, 5-layer architecture, DAMA-DMBOK governance implemented as executable tests.
Convention-aware integrity auditor for SCD Type 2 dimensions. 16 SQL checks covering overlaps, gaps, current-flag drift, sentinel drift and point-in-time referential breaks, with findings ranked by the fact rows and revenue at risk and idempotent repair SQL in DuckDB and Snowflake via sqlglot. 1.28M version rows audited in 3.95s.
Modern data stack reference: dbt + BigQuery + Airflow (Cloud Composer) with medallion layering, SCD2 snapshots, exposures, freshness SLAs, and 45× cost reduction via partition + cluster + incremental tuning.
Batch medallion lakehouse: raw -> bronze/silver Delta (delta-rs) -> dbt gold star schema with a full SCD2 customer dimension and point-in-time fact resolution. Great Expectations guards each layer and a reconciler proves gold equals silver in dollars. Ships Spark jobs, an Airflow DAG, and Terraform.
Pipeline ELT em Python e dbt sobre as APIs do Banco Central: projeções do Focus contra o realizado do SGS, com SCD Tipo 2 sobre revisões de série.
End-to-end real-time data pipeline using Apache NiFi, AWS (EC2/S3), and Snowflake — auto-ingests streaming customer data via Snowpipe and tracks full change history using Snowflake Streams and Tasks (SCD Type 2).
Read-only temporal-integrity validator for SCD Type 2 dimensions. Runs seven DuckDB window-function checks (overlaps, gaps, multiple current rows, and more) and reports each violation's blast radius in keys, fact rows, and dollars.
To associate your repository with the scd2 topic, visit your repo's landing page and select "manage topics."