Bucketing splits a table's rows into a fixed number of files by hashing a column. Partitioning splits a table into directories, one per value of a column. They solve different problems: partitioning lets queries skip whole directories, and bucketing lets joins and aggregations skip the shuffle.
How each looks on disk
Partitioned by order_date: Bucketed by customer_id, 4 buckets:
order_date=2025-03-01/ part-0 (hash(customer_id) % 4 = 0)
order_date=2025-03-02/ part-1 (hash(customer_id) % 4 = 1)
order_date=2025-03-03/ part-2 (hash(customer_id) % 4 = 2)
part-3 (hash(customer_id) % 4 = 3)Writing a bucketed table
(df.write
.bucketBy(64, "customer_id")
.sortBy("customer_id")
.mode("overwrite")
.saveAsTable("sales.orders_bucketed"))It must be saveAsTable, because the bucket information is stored in the metastore. A plain .parquet(path) write cannot record it.
Why it speeds up joins
If two tables are bucketed on the same column into the same number of buckets, then rows with the same key are already in the same bucket number on both sides. Spark can join bucket 5 with bucket 5 without moving anything. It saves the shuffle (and the sort, if you also sorted by the key). The benefit is large if you join the same big tables again and again.
When to use which
- Partitioning: low-cardinality columns that queries filter by, such as date or country. High-cardinality columns create too many directories.
- Bucketing: high-cardinality join or group keys, such as
customer_id, used repeatedly.
Practical limits
- Both tables need the same bucket count (or a multiple, with some settings) and the same key.
- Writing is more expensive and you must keep the layout consistent. A new writer that forgets
bucketBybreaks the benefit. - Many modern table formats do not use Spark-style bucketing. Delta Lake does not support
bucketBy, and relies on partitioning plus clustering (Z-order or liquid clustering). Iceberg has a bucket transform that you can use in its partition spec. Because of this, many lakehouse teams use AQE and broadcast joins rather than bucketing.