Clustering organizes rows inside a table so that rows with similar values for chosen columns are stored near each other. That tightens file-level statistics and improves data skipping for filters on those columns, without forcing a rigid Hive partition tree.
Partitioning vs clustering
Partitioning: separate directories/files by key (dt=2024-01-01/...) Clustering: sort/co-locate within the table layout by keys (user_id, ...)
Variants you may hear
- Explicit sort/cluster-by on write
- Delta Z-Order clustering via
OPTIMIZE - Delta liquid clustering (cluster keys maintained more automatically over writes)
- Warehouse
CLUSTER BY/ automatic micro-partition clustering (Snowflake-style mental model)
Why teams cluster
- High-cardinality filter columns that are bad partition keys
- Avoid partition explosion while still accelerating common predicates
- Improve join locality in some engines when keys align
Costs
- Rewrites or ongoing maintenance CPU
- Wrong cluster keys waste effort
- Still need reasonable file sizes (clustering does not fix 100k tiny files alone)
Interview tip: "Partitions answer coarse filters (date); clustering answers selective filters inside those slices." Mention Z-order / liquid clustering if the stack is Delta.