This project is a data transformation and analytics pipeline built using dbt (data build tool) and Snowflake. It uses the MovieLens dataset to transform raw movie, rating, tag, genome tag, genome score, and link data into structured dimension and fact tables.
The project demonstrates core dbt concepts including:
- Sources
- Staging models
- Dimension and fact models
- Mart models
- Seeds
- Macros
- Snapshots
- Generic and custom tests
- Incremental models
- Ephemeral models
- dbt packages
- Model documentation
- Jinja templating
- Model dependencies using
ref()
MovieLens Raw Data
|
v
Raw Schema / Tables
|
v
Staging
|
+------------------+------------------+
| | |
v v v
Dimension Fact Models Snapshots
Models
|
v
Marts
|
v
Analytics / BI
Detailed flow:
RAW_MOVIES
|
v
src_movies
|
v
dim_movies
|
+----------------------+
| |
v v
dim_movies_with_tags mart_movie_releases
|
v
ep_movies_with_tags
RAW_RATINGS
|
v
src_ratings
|
v
fct_ratings
|
v
mart_movie_releases
RAW_TAGS
|
+------------------+
| |
v v
src_tags snap_tags
|
v
dim_users
RAW_GENOME_TAGS
|
v
src_genome_tags
|
v
dim_genome_tags
|
v
dim_movies_with_tags
RAW_GENOME_SCORES
|
v
src_genome_score
|
v
fct_genome_scores
|
v
dim_movies_with_tags
- dbt
- Snowflake
- SQL
- Jinja
- dbt-utils
- Git
netflix/
│
├── dbt_project.yml
├── packages.yml
│
├── models/
│ ├── staging/
│ │ ├── src_genome_score.sql
│ │ ├── src_genome_tags.sql
│ │ ├── src_links.sql
│ │ ├── src_movies.sql
│ │ ├── src_ratings.sql
│ │ └── src_tags.sql
│ │
│ ├── dim/
│ │ ├── dim_genome_tags.sql
│ │ ├── dim_movies.sql
│ │ ├── dim_movies_with_tags.sql
│ │ └── dim_users.sql
│ │
│ ├── fct/
│ │ ├── ep_movies_with_tags.sql
│ │ ├── fct_genome_scores.sql
│ │ └── fct_ratings.sql
│ │
│ └── mart/
│ | └── mart_movie_releases.sql
| ├── schema.yml
| └── sources.yml
│
├── seeds/
│ └── seed_movie_release_dates.csv
│
├── snapshots/
│ └── snap_tags.sql
│
├── macros/
│ └── no_nulls_in_columns.sql
│
├── tests/
│ └── relevance_score_test.sql
│
└── analyses/
└── movie_analysis.sql
The project uses raw MovieLens data stored in Snowflake.
The raw tables include:
RAW_MOVIES
RAW_RATINGS
RAW_TAGS
RAW_GENOME_TAGS
RAW_GENOME_SCORES
RAW_LINKS
These tables are declared as dbt sources in sources.yml.
Example:
sources:
- name: netflix
schema: raw
tables:
- name: r_movies
identifier: raw_moviesThe name is the logical name used by dbt, while identifier represents the actual physical table name.
Sources can be referenced using:
{{ source('netflix', 'r_movies') }}The staging layer performs basic transformations on raw data.
Typical operations include:
- Renaming columns
- Standardizing column names
- Converting data types
- Removing unnecessary columns
- Basic data cleaning
For example:
SELECT
movieId AS movie_id,
title,
genres
FROM ...The staging layer converts camelCase column names such as:
movieId
userId
tagId
into standardized names:
movie_id
user_id
tag_id
Timestamp fields are also converted where required:
TO_TIMESTAMP_LTZ(timestamp) AS rating_timestampThe project contains the following dimension models:
dim_movies
dim_users
dim_genome_tags
dim_movies_with_tags
Cleans and standardizes movie metadata.
It performs transformations such as:
INITCAP(TRIM(title)) AS movie_titleand:
SPLIT(genres, '|') AS genre_arrayThis provides both the original genre string and an array representation.
Creates a unique list of users appearing in both ratings and tags.
It uses:
UNIONto combine users from the two datasets.
Cleans genome tag names using:
INITCAP(TRIM(tag))Combines movie information with genome tags and relevance scores.
It uses LEFT JOIN operations between:
dim_movies
fct_genome_scores
dim_genome_tags
This model is configured as:
{{ config(materialized='ephemeral') }}Therefore, it is not stored as a physical table or view. Its SQL is incorporated into downstream models during compilation.
The project contains:
fct_ratings
fct_genome_scores
Stores user movie ratings.
It is configured as an incremental model:
{{ config(
materialized='incremental',
on_schema_change='fail'
) }}The model filters invalid ratings:
WHERE rating IS NOT NULLDuring incremental runs, only records newer than the latest timestamp already present in the target table are processed:
{% if is_incremental() %}
AND rating_timestamp > (
SELECT MAX(rating_timestamp)
FROM {{ this }}
)
{% endif %}This avoids rebuilding the entire dataset on every run.
Stores the relevance of genome tags for movies.
It:
- Removes non-positive relevance scores
- Rounds scores to four decimal places
ROUND(relevance, 4) AS relevance_scoreand:
WHERE relevance > 0The project contains:
mart_movie_releases
This is a business-oriented table intended for downstream analytics.
It combines:
fct_ratings
+
seed_movie_release_dates
The model uses a LEFT JOIN to retain rating records even when a release date is unavailable.
It also creates a derived field:
CASE
WHEN release_date IS NULL THEN 'unknown'
ELSE 'known'
END AS release_info_availableThe model is materialized as a table.
The project contains:
seed_movie_release_dates.csv
Example:
movie_id,release_date
1,1995-10-20
2,1995-10-21
3,1995-10-25Seeds are useful for small, relatively static datasets that can be maintained alongside the dbt project.
Run:
dbt seedto load the CSV into the warehouse.
The resulting relation can be referenced with:
{{ ref('seed_movie_release_dates') }}The project contains:
snapshots/snap_tags.sql
The snapshot is used to track historical changes to tag data.
It uses:
strategy='timestamp'with:
updated_at='tag_timestamp'The logical unique key is:
user_id + movie_id + tag
Configuration:
{{ config(
target_schema='snapshots',
unique_key=['user_id','movie_id','tag'],
strategy='timestamp',
updated_at='tag_timestamp',
invalidate_hard_deletes=True
) }}Snapshots allow historical versions of records to be preserved rather than only retaining the current state.
The project contains a custom macro:
macros/no_nulls_in_columns.sql
The macro dynamically examines the columns of a relation using:
adapter.get_columns_in_relation(model)It is intended to generate a query that checks whether any column contains NULL values.
Macros allow reusable SQL/Jinja logic to be written once and used across multiple models or tests.
The project uses the dbt-utils package.
packages.yml contains:
packages:
- package: dbt-labs/dbt_utils
version: 1.3.0Install project dependencies using:
dbt depsThe project uses the package's surrogate-key macro:
{{ dbt_utils.generate_surrogate_key(
['user_id','movie_id','tag']
) }}This generates a deterministic surrogate key from multiple columns.
The project uses both generic and custom dbt tests.
Generic tests are defined in schema.yml.
Examples include:
tests:
- not_nulland:
tests:
- relationships:
to: ref('dim_movies')
field: movie_idThe relationship test ensures that every movie_id in fct_ratings exists in dim_movies.
Conceptually:
fct_ratings.movie_id
|
v
dim_movies.movie_id
Custom SQL tests are stored in:
tests/
The relevance score test checks that invalid relevance scores are not present.
A dbt data test generally passes when its query returns zero rows.
Run tests using:
dbt testThe project uses schema.yml to document models and columns.
Example:
- name: dim_movies
description: Dimension table for cleansed movie metadataColumns can also be documented:
- name: movie_id
description: Primary key of the movieThe same YAML file can contain tests.
This allows documentation and data-quality rules to be maintained alongside the models.
The project contains:
analyses/movie_analysis.sql
The analysis calculates movie-level rating statistics.
It calculates:
average_rating
total_ratings
and only considers movies with more than 100 ratings:
HAVING COUNT(*) > 100The results are ordered by average rating and limited to 20 movies.
Analyses are useful for analytical SQL that does not need to be materialized as a standard dbt model.
The project demonstrates multiple dbt materializations.
The project default is:
+materialized: viewViews are virtual database objects whose underlying query is executed when accessed.
Dimension and fact folders are configured as tables:
dim:
+materialized: table
fct:
+materialized: tableIndividual models can also override the default using:
{{ config(materialized='table') }}fct_ratings uses:
{{ config(materialized='incremental') }}Only new data is processed according to the incremental filter.
dim_movies_with_tags uses:
{{ config(materialized='ephemeral') }}The model is not physically created in the warehouse. Its SQL is injected into downstream models.
Used to reference external/raw tables:
{{ source('netflix', 'r_movies') }}Used to reference another dbt relation:
{{ ref('src_movies') }}It also creates a dependency in the dbt DAG.
Used to configure a model:
{{ config(materialized='table') }}Checks whether the current model is being executed as an incremental update:
{% if is_incremental() %}
...
{% endif %}References the current model's target relation:
{{ this }}The dbt DAG is automatically created from dependencies such as ref() and source().
For example:
source: netflix.r_movies
|
v
src_movies
|
v
dim_movies
|
v
dim_movies_with_tags
dbt uses these dependencies to determine the correct execution order.
You do not need to manually execute every model in dependency order.
Install packages:
dbt depsLoad seeds:
dbt seedRun models:
dbt runRun tests:
dbt testRun snapshots:
dbt snapshotCompile SQL without executing it:
dbt compileBuild models and other supported resources according to their dependencies:
dbt buildRemove generated directories:
dbt cleanCheck the project configuration and connection:
dbt debugGenerate project documentation:
dbt docs generateServe the documentation locally:
dbt docs serveThis project demonstrates the following dbt concepts:
| Concept | Demonstrated In |
|---|---|
| Project configuration | dbt_project.yml |
| Source definitions | sources.yml |
| Staging models | models/staging/ |
| Dimension models | models/dim/ |
| Fact models | models/fct/ |
| Mart models | models/mart/ |
| Seeds | seeds/ |
| Snapshots | snapshots/ |
| Custom macros | macros/ |
| Generic tests | schema.yml |
| Custom SQL tests | tests/ |
| Documentation | schema.yml |
| Incremental models | fct_ratings.sql |
| Ephemeral models | dim_movies_with_tags.sql |
| External packages | packages.yml |
source() |
Staging models |
ref() |
Downstream models |
| Jinja | Models, macros and snapshots |
| CTEs | Most transformation models |
| SQL joins | Dimension and mart models |
| SQL aggregation | Analysis |
| Surrogate keys | Snapshot |
The project follows a dimensional modeling approach.
Dimension tables:
dim_movies
dim_users
dim_genome_tags
Fact tables:
fct_ratings
fct_genome_scores
The general relationship is:
dim_movies
▲
│
movie_id
│
fct_ratings
│
user_id
│
▼
dim_users
For genome scores:
dim_movies
▲
│
movie_id
│
fct_genome_scores
│
tag_id
│
▼
dim_genome_tags
This provides a structured foundation for analytics and BI applications.
The primary goal of this project is to demonstrate how raw MovieLens data can be transformed into a maintainable analytics-ready data warehouse using dbt.
The project applies:
Raw Data
|
v
Source Definitions
|
v
Staging
|
v
Dimensions + Facts
|
v
Marts
|
v
Analytics
while incorporating data quality testing, documentation, incremental processing, historical tracking, reusable macros, and dependency management.
Clone the repository and navigate to the project:
git clone <repository-url>
cd netflixInstall dbt package dependencies:
dbt depsVerify the dbt configuration:
dbt debugLoad seed data:
dbt seedRun the project:
dbt runRun tests:
dbt testOr build the project and execute supported resources according to their dependencies:
dbt buildGenerate documentation:
dbt docs generateStart the documentation server:
dbt docs servedbt_project.ymlcontrols project-level configuration.sources.ymldeclares raw warehouse tables as dbt sources.ref()should be used when referencing another dbt model.source()should be used when referencing declared external/raw tables.fct_ratingsuses incremental processing.dim_movies_with_tagsis ephemeral and is not stored as a physical warehouse relation.seed_movie_release_dates.csvprovides manually maintained movie release-date information.snap_tags.sqltracks historical changes to tag records.schema.ymlprovides model documentation and data-quality tests.dbt_utilsprovides reusable macros used by the project.