Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Liquid clustering

File formats & storage · Table Formats in Depth

Liquid clustering

Mediumfile-formats-38
delta-lakeliquid-clusteringpartitioningz-orderdatabricks

Question

What is liquid clustering in Delta Lake, and why does it replace partitioning + Z-order?

Solution

Liquid clustering is a modern layout optimization technique in Delta Lake that organizes data dynamically along specified clustering keys without relying on static Hive directory partitions or rigid multi-dimensional Z-order algorithms. It replaces the traditional combination of folder partitioning and periodic OPTIMIZE ZORDER BY by clustering data incrementally during writes and maintenance jobs, while allowing data teams to change clustering columns at any time without rewriting existing table files. Because it adapts layout flexibly, it prevents small-file fragmentation and metadata explosion on high-cardinality columns.

Flexible clustering without static directories

Legacy table organization forced engineering teams into painful architectural compromises:

  • Hive directory partitioning required choosing low-cardinality columns (like date). If you partitioned by user_id or device_id, you created millions of folders containing tiny 10 KB files, crippling query planning and metastores.
  • Z-order improved multi-column skipping within partitions, but it was non-incremental: running OPTIMIZE with Z-order required reading and rewriting every data file in the partition, consuming heavy cluster compute.
  • Partitioning columns were permanently locked; changing partition columns required creating a new table and rewriting all historical data.

Incremental clustering operations

Liquid clustering solves these challenges through dynamic layout structures:

  • Changeable keys: You can redefine clustering keys with a simple ALTER TABLE statement. New writes immediately use the updated clustering keys without triggering full historical rewrites.
  • Incremental clustering: During regular ingestion and scheduled OPTIMIZE runs, liquid clustering clusters only newly ingested or unclustered files, avoiding expensive full-table rewrites.
  • High-cardinality support: Columns like customer_id, device_id, or fine-grained timestamps can be used as clustering keys without generating physical directory trees.

When to transition from Hive partitioning

Liquid clustering is the standard recommendation for all new Delta tables on modern Databricks runtimes. However, remember that liquid clustering and traditional Hive partitioning are mutually exclusive; a Delta table cannot use both simultaneously. If migrating an existing partitioned table, you must rewrite the table into an unpartitioned structure configured with CLUSTER BY.

PreviousNext