Partitioning splits a large table into chunks (often by date) so queries and pipeline jobs can scan only relevant pieces.
fct_events partitioned by event_date query for 2026-09-05 -> read only that partition, not 3 years of files
Benefits
- Faster queries with partition pruning
- Cheaper loads (replace one day)
- Easier retention deletes (drop old partitions)
- Safer backfills (rebuild one partition)
Common partition keys
event_date, ingest_date, country (less common alone)
Gotchas
Too many tiny partitions (over-partitioning) hurts planning and file sizes. Pick a grain that matches query filters.
Interview tip: Connect partitioning to pruning, idempotent daily loads, and cost.