Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Partitioning in file storage

File formats & storage · Formats & Table Formats

Partitioning in file storage

Easyformat-10
partitioningHive styleobject storagelayout

Question

How does partitioning work for datasets stored as files?

Solution

Partitioning in file storage means organizing files into directories (or table-format partitions) keyed by column values so queries can skip entire folders.

Classic Hive-style layout:

/lake/events/
  dt=2024-01-01/region=us/...
  dt=2024-01-01/region=eu/...
  dt=2024-01-02/region=us/...

A query WHERE dt = '2024-01-02' only lists and reads the matching dt= path. That is partition pruning.

What partitioning gives you

  • Smaller listing and scan scope
  • Parallelism aligned with day/region slices
  • Cheaper lifecycle policies (expire old dt= folders)

Costs and traps

  • Too many partition values → tiny folders and the small-files problem
  • Partitioning on high-cardinality keys (user_id) creates millions of directories
  • Changing partition strategy later is painful without Iceberg partition evolution / rewrites

Interview tip: "Partitions are a physical layout optimization for filters you almost always apply (often date)." Mention table formats still use partitions or clustering under the hood.

PreviousNext