Skip to content
Great engineers know 3 min read · Iceberg 1 practice problem ↓

Hidden partitioning and partition evolution

Iceberg partition transforms, and changing a table layout without rewriting its data.

You will learn

  • How Iceberg partition transforms work
  • Why users never filter on partition columns directly
  • How to change a table's partitioning without rewriting data
  • How planning handles data written with different layouts

Read first

Comfortable with these? Read on.

TL;DR Iceberg partitions by transforms of real columns, such as day(ts) or bucket(16, user_id). Queries filter on the original column and still prune. Because partitioning lives in metadata, you can change it later (partition evolution) and old data keeps its old layout.

The problem with Hive-style partitions

In Hive-style tables, the partition value is a separate column that writers must fill and readers must filter on. To partition events by day you add an event_date column derived from ts. Then a query on ts alone (WHERE ts >= '2025-03-01 10:00') does not prune, because the engine does not know event_date comes from ts. Writers that compute event_date in a different time zone silently put rows in the wrong partition.

Partition transforms

TransformPartitions byExample
identity(col)The value itselfcountry
year, month, day, hourTime granularity of a date or timestampday(ts)
bucket(N, col)Hash of the value into N bucketsbucket(16, user_id)
truncate(W, col)Value truncated to width W (strings, numbers)truncate(4, zip_code)
Spark SQL
CREATE TABLE db.events (
  event_id BIGINT, user_id BIGINT, ts TIMESTAMP, payload STRING)
USING iceberg
PARTITIONED BY (day(ts), bucket(16, user_id));

-- Users filter on real columns; Iceberg derives the partition filter
SELECT * FROM db.events
WHERE ts >= TIMESTAMP '2025-03-01 10:00:00' AND user_id = 42;

Writers do not provide partition values; Iceberg computes them from the data, so they are always consistent. Readers filter on ts and user_id; Iceberg translates those predicates into partition predicates (the matching days, and the bucket of 42).

Partition evolution

2023 · month(ts)

The table starts small, partitioned by month(ts).

2025 · grown

Volume grows 50 times; monthly partitions are now huge. You switch to daily partitions with a metadata-only change.

The command

ALTER TABLE db.events REPLACE PARTITION FIELD month(ts) WITH day(ts);

After · two specs

Files written before the change keep the monthly spec; new files use the daily spec. Each manifest records which spec its files use. No data is rewritten.

Queries · still prune

A query for March 14 prunes old data at month level and new data at day level, planning each spec separately. Users notice nothing, except faster queries on recent data.

Other evolution statements: ALTER TABLE ... ADD PARTITION FIELD bucket(16, id) and DROP PARTITION FIELD. If you want old data in the new layout too, rewrite it with rewrite_data_files at your own pace.

Tip: start coarse. Because evolution is cheap, a table can begin with monthly or no partitioning and move to finer partitions when volume justifies it.

Delta Lake approaches the same problem differently: generated columns can derive a partition column from another column (with partition filters derived automatically for some expressions), and liquid clustering replaces fixed partitioning with clustering keys that can change.

Common mistakes

Adding a separate date column out of Hive habit

Use day(ts) and filter on ts.

Using identity on high-cardinality columns

Use bucket(N, col) instead.

Assuming evolution rewrites old data

Old files keep the old spec until rewritten.

Key takeaways

  • Iceberg partitions by transforms: identity, year/month/day/hour, bucket, truncate.
  • Queries filter on real columns and still prune.
  • Partition evolution changes the layout for new data with no rewrite.
  • Start coarse and evolve as data grows.

Check yourself

3 questions

1. Which filter prunes a table partitioned by day(ts)?

Show the answer

WHERE ts >= TIMESTAMP .... Iceberg derives partition predicates from filters on the source column.

2. What happens to existing files after REPLACE PARTITION FIELD month(ts) WITH day(ts)?

Show the answer

They keep the old spec; new files use the new one. Partition evolution is metadata-only.

3. Which transform suits a high-cardinality user_id?

Show the answer

bucket(16, user_id). Bucketing hashes values into a fixed number of partitions.

Practice it

Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.

Solve: Iceberg Hidden Partition Layout →

Go deeper

Primary sources: Iceberg: partitioning · Iceberg: partition evolution