Liquid clustering
Incremental clustering in Delta Lake that replaces fixed partitions and Z-order.
On this page
Show code in
Every code block on the page follows this.
You will learn
- What liquid clustering is and the problems with partitions and Z-order it solves
- How to create, change and maintain clustering keys
- How clustering happens incrementally
- When to use it and when not
Read first
- Compaction, OPTIMIZE and Z-order · 3 min read
- Partitioning done right · 4 min read
Comfortable with these? Read on.
CLUSTER BY. OPTIMIZE clusters data incrementally, only rewriting what is not yet clustered, and you can change the keys later without rewriting existing data.What it fixes
| Partitioning | Z-order | Liquid clustering |
|---|---|---|
| Fixed at creation; changing it means rewriting the table | Must be re-run, and re-running rewrites data already ordered | Keys can change; new data uses new keys |
| High-cardinality columns create tiny files | Works on high cardinality but is not incremental | Works on high cardinality, incremental |
| Skewed partitions are common | Expensive on large tables | Balances file sizes automatically |
Using it
(df.writeTo("events") .using("delta") .clusterBy("event_date", "customer_id") .create())
CREATE TABLE events (event_id BIGINT, customer_id BIGINT, event_date DATE, payload STRING) USING delta CLUSTER BY (event_date, customer_id); -- Cluster newly written data (incremental) OPTIMIZE events; -- Change the keys later: existing data is not rewritten immediately ALTER TABLE events CLUSTER BY (customer_id); -- Remove clustering ALTER TABLE events CLUSTER BY NONE;
- Up to 4 clustering columns; choose columns used in filters, high or low cardinality.
- Clustering cannot be combined with
PARTITIONED BYorZORDER BYon the same table. - Existing partitioned tables cannot be converted in place in all versions; check your Delta release, or create a new clustered table.
How it works
Data is laid out along a space-filling curve (a Hilbert curve) over the clustering columns, which preserves locality in several dimensions better than Z-order. Writes land unclustered or partially clustered; OPTIMIZE finds files that are not yet well clustered and rewrites only those, grouping them into balanced, well-sized files. Already clustered data is left alone, which is what makes it incremental and cheap to run frequently.
When to use it
- New tables where the query pattern may change.
- Tables filtered by high-cardinality columns (ids) or by several columns.
- Tables whose partitions would be skewed or tiny.
- Not needed for small tables that are fully scanned anyway; and check that every engine reading the table supports the clustering table feature.
Common mistakes
Expecting data to be clustered right after writing
Combining with partitioning or Z-order
Using too many clustering columns
Key takeaways
- Liquid clustering replaces partitioning and Z-order with CLUSTER BY keys.
- OPTIMIZE clusters only data that is not yet clustered.
- Keys can change without rewriting existing data.
- Up to 4 keys; not combinable with partitions or Z-order.
Check yourself
3 questions1. What does OPTIMIZE do on a liquid-clustered table?
Show the answer
Clusters only files that are not yet well clustered. Clustering is incremental.
2. What happens to existing data after ALTER TABLE ... CLUSTER BY (new_col)?
Show the answer
It stays; new clustering applies as data is optimized. Changing keys is a metadata change.
3. Can a table use both PARTITIONED BY and CLUSTER BY?
Show the answer
No. Liquid clustering replaces partitioning on that table.
Go deeper
Compaction, OPTIMIZE and Z-orderLakehouse
Partitioning done rightLakehouse
Hidden partitioning and partition evolution
Primary sources: Delta: liquid clustering