Skip to content
Great engineers know 3 min read · Delta

Liquid clustering

Incremental clustering in Delta Lake that replaces fixed partitions and Z-order.

On this page

Show code in

Every code block on the page follows this.

You will learn

  • What liquid clustering is and the problems with partitions and Z-order it solves
  • How to create, change and maintain clustering keys
  • How clustering happens incrementally
  • When to use it and when not

Read first

Comfortable with these? Read on.

TL;DR Liquid clustering replaces partitioning and Z-order in Delta Lake with clustering keys declared by CLUSTER BY. OPTIMIZE clusters data incrementally, only rewriting what is not yet clustered, and you can change the keys later without rewriting existing data.

What it fixes

PartitioningZ-orderLiquid clustering
Fixed at creation; changing it means rewriting the tableMust be re-run, and re-running rewrites data already orderedKeys can change; new data uses new keys
High-cardinality columns create tiny filesWorks on high cardinality but is not incrementalWorks on high cardinality, incremental
Skewed partitions are commonExpensive on large tablesBalances file sizes automatically

Using it

PySparkSpark SQL
(df.writeTo("events")
    .using("delta")
    .clusterBy("event_date", "customer_id")
    .create())
CREATE TABLE events (event_id BIGINT, customer_id BIGINT, event_date DATE, payload STRING)
USING delta
CLUSTER BY (event_date, customer_id);

-- Cluster newly written data (incremental)
OPTIMIZE events;

-- Change the keys later: existing data is not rewritten immediately
ALTER TABLE events CLUSTER BY (customer_id);

-- Remove clustering
ALTER TABLE events CLUSTER BY NONE;
  • Up to 4 clustering columns; choose columns used in filters, high or low cardinality.
  • Clustering cannot be combined with PARTITIONED BY or ZORDER BY on the same table.
  • Existing partitioned tables cannot be converted in place in all versions; check your Delta release, or create a new clustered table.

How it works

Data is laid out along a space-filling curve (a Hilbert curve) over the clustering columns, which preserves locality in several dimensions better than Z-order. Writes land unclustered or partially clustered; OPTIMIZE finds files that are not yet well clustered and rewrites only those, grouping them into balanced, well-sized files. Already clustered data is left alone, which is what makes it incremental and cheap to run frequently.

Note: on Databricks, liquid clustering is the recommended layout for new tables, with options such as automatic key selection and clustering on write. Open-source Delta Lake supports the core feature; check your release notes for which extras are available.

When to use it

  • New tables where the query pattern may change.
  • Tables filtered by high-cardinality columns (ids) or by several columns.
  • Tables whose partitions would be skewed or tiny.
  • Not needed for small tables that are fully scanned anyway; and check that every engine reading the table supports the clustering table feature.

Common mistakes

Expecting data to be clustered right after writing

Run OPTIMIZE (or enable clustering on write where available).

Combining with partitioning or Z-order

Not allowed on the same table.

Using too many clustering columns

Each additional column weakens clustering; maximum 4.

Key takeaways

  • Liquid clustering replaces partitioning and Z-order with CLUSTER BY keys.
  • OPTIMIZE clusters only data that is not yet clustered.
  • Keys can change without rewriting existing data.
  • Up to 4 keys; not combinable with partitions or Z-order.

Check yourself

3 questions

1. What does OPTIMIZE do on a liquid-clustered table?

Show the answer

Clusters only files that are not yet well clustered. Clustering is incremental.

2. What happens to existing data after ALTER TABLE ... CLUSTER BY (new_col)?

Show the answer

It stays; new clustering applies as data is optimized. Changing keys is a metadata change.

3. Can a table use both PARTITIONED BY and CLUSTER BY?

Show the answer

No. Liquid clustering replaces partitioning on that table.

Go deeper

Primary sources: Delta: liquid clustering