Day 37 of 45 · Pro lesson
The curriculum for this day stays visible. The lesson, code, and workspace unlock with Pro.
Teaches: Medallion curation: Bronze raw to Silver cleaned to Gold aggregates, Windowed deduplication by primary key and updated_at timestamp, Data contract casting & null substitution policies, Broadcast dimension enrichment, Gold daily revenue & customer lifetime value fact tables, Atomic S3 partition swaps
Build: Platform Stage 2: PySpark Silver deduplication & validation engine → Gold metric aggregations → broadcast join enricher → atomic lakehouse writer.
Always deduplicate Silver data using row_number() over (partition by id order by updated_at desc) before performing any join operations.
Already on Pro? Go to your account. Need a pass? See pricing.