Partitioning in file storage means organizing files into directories (or table-format partitions) keyed by column values so queries can skip entire folders.
Classic Hive-style layout:
/lake/events/ dt=2024-01-01/region=us/... dt=2024-01-01/region=eu/... dt=2024-01-02/region=us/...
A query WHERE dt = '2024-01-02' only lists and reads the matching dt= path. That is partition pruning.
What partitioning gives you
- Smaller listing and scan scope
- Parallelism aligned with day/region slices
- Cheaper lifecycle policies (expire old
dt=folders)
Costs and traps
- Too many partition values → tiny folders and the small-files problem
- Partitioning on high-cardinality keys (
user_id) creates millions of directories - Changing partition strategy later is painful without Iceberg partition evolution / rewrites
Interview tip: "Partitions are a physical layout optimization for filters you almost always apply (often date)." Mention table formats still use partitions or clustering under the hood.