Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is clustering?

File formats & storage · Extra High-Value

What is clustering?

Hardformat-19
clusteringliquid clusteringZ-orderlayoutlakehouse

Question

What is clustering (or liquid clustering) for lakehouse tables?

Solution

Clustering organizes rows inside a table so that rows with similar values for chosen columns are stored near each other. That tightens file-level statistics and improves data skipping for filters on those columns, without forcing a rigid Hive partition tree.

Partitioning vs clustering

Partitioning: separate directories/files by key (dt=2024-01-01/...)
Clustering:  sort/co-locate within the table layout by keys (user_id, ...)

Variants you may hear

  • Explicit sort/cluster-by on write
  • Delta Z-Order clustering via OPTIMIZE
  • Delta liquid clustering (cluster keys maintained more automatically over writes)
  • Warehouse CLUSTER BY / automatic micro-partition clustering (Snowflake-style mental model)

Why teams cluster

  • High-cardinality filter columns that are bad partition keys
  • Avoid partition explosion while still accelerating common predicates
  • Improve join locality in some engines when keys align

Costs

  • Rewrites or ongoing maintenance CPU
  • Wrong cluster keys waste effort
  • Still need reasonable file sizes (clustering does not fix 100k tiny files alone)

Interview tip: "Partitions answer coarse filters (date); clustering answers selective filters inside those slices." Mention Z-order / liquid clustering if the stack is Delta.

PreviousNext