OPTIMIZE rewrites many small files into fewer large ones. ZORDER BY sorts data within those files so related values sit together and queries can skip files. VACUUM deletes old files that are no longer part of the table. Together they keep a Delta table fast and its storage under control.
OPTIMIZE
OPTIMIZE sales.orders;
Streaming and frequent small writes leave thousands of small files. Reading them is slow because each file has an overhead for listing, opening and planning. OPTIMIZE compacts them toward a target size of about 1 GB per file. It does not change the data. Old files stay on storage until they are vacuumed, and readers keep working during the rewrite.
ZORDER BY
OPTIMIZE sales.orders ZORDER BY (customer_id, product_id);
Delta keeps min and max statistics per file for the first columns, which allows data skipping. Z-ordering arranges rows so that rows with close values of the chosen columns end up in the same files, across several columns together, so a filter on customer_id reads few files. It rewrites the data, so it costs compute, and its benefit fades as columns increase. Choose columns that you often filter on.
VACUUM
VACUUM sales.orders; -- removes unreferenced files older than the retention period
After OPTIMIZE, updates and deletes, old files remain for time travel and for readers still using them. VACUUM deletes files not referenced by the table and older than the retention threshold, 7 days by default. After that, you cannot time travel to versions that needed those files. Do not set the retention too low: a long-running query or a stream reading an old version can fail if its files are removed.
Liquid clustering
For new tables, Databricks recommends liquid clustering instead of partitioning plus Z-order:
CREATE TABLE sales.orders (...) CLUSTER BY (customer_id, order_date);
You can change the clustering keys later without rewriting the whole table, and it handles incremental clustering. It replaces the need to pick a fixed partition scheme and rerun Z-order.
Summary
Compact with OPTIMIZE, cluster for skipping, clean up with VACUUM, and use predictive optimization (next question) to automate it.