Hidden partitioning and partition evolution
Iceberg partition transforms, and changing a table layout without rewriting its data.
On this page
Show code in
Every code block on the page follows this.
You will learn
- How Iceberg partition transforms work
- Why users never filter on partition columns directly
- How to change a table's partitioning without rewriting data
- How planning handles data written with different layouts
Read first
- Iceberg metadata tree · 4 min read
- Partitioning done right · 4 min read
Comfortable with these? Read on.
day(ts) or bucket(16, user_id). Queries filter on the original column and still prune. Because partitioning lives in metadata, you can change it later (partition evolution) and old data keeps its old layout.The problem with Hive-style partitions
In Hive-style tables, the partition value is a separate column that writers must fill and readers must filter on. To partition events by day you add an event_date column derived from ts. Then a query on ts alone (WHERE ts >= '2025-03-01 10:00') does not prune, because the engine does not know event_date comes from ts. Writers that compute event_date in a different time zone silently put rows in the wrong partition.
Partition transforms
| Transform | Partitions by | Example |
|---|---|---|
identity(col) | The value itself | country |
year, month, day, hour | Time granularity of a date or timestamp | day(ts) |
bucket(N, col) | Hash of the value into N buckets | bucket(16, user_id) |
truncate(W, col) | Value truncated to width W (strings, numbers) | truncate(4, zip_code) |
CREATE TABLE db.events ( event_id BIGINT, user_id BIGINT, ts TIMESTAMP, payload STRING) USING iceberg PARTITIONED BY (day(ts), bucket(16, user_id)); -- Users filter on real columns; Iceberg derives the partition filter SELECT * FROM db.events WHERE ts >= TIMESTAMP '2025-03-01 10:00:00' AND user_id = 42;
Writers do not provide partition values; Iceberg computes them from the data, so they are always consistent. Readers filter on ts and user_id; Iceberg translates those predicates into partition predicates (the matching days, and the bucket of 42).
Partition evolution
2023 · month(ts)
month(ts).2025 · grown
The command
ALTER TABLE db.events REPLACE PARTITION FIELD month(ts) WITH day(ts);
After · two specs
Queries · still prune
Other evolution statements: ALTER TABLE ... ADD PARTITION FIELD bucket(16, id) and DROP PARTITION FIELD. If you want old data in the new layout too, rewrite it with rewrite_data_files at your own pace.
Delta Lake approaches the same problem differently: generated columns can derive a partition column from another column (with partition filters derived automatically for some expressions), and liquid clustering replaces fixed partitioning with clustering keys that can change.
Common mistakes
Adding a separate date column out of Hive habit
Using identity on high-cardinality columns
Assuming evolution rewrites old data
Key takeaways
- Iceberg partitions by transforms: identity, year/month/day/hour, bucket, truncate.
- Queries filter on real columns and still prune.
- Partition evolution changes the layout for new data with no rewrite.
- Start coarse and evolve as data grows.
Check yourself
3 questions1. Which filter prunes a table partitioned by day(ts)?
Show the answer
WHERE ts >= TIMESTAMP .... Iceberg derives partition predicates from filters on the source column.
2. What happens to existing files after REPLACE PARTITION FIELD month(ts) WITH day(ts)?
Show the answer
They keep the old spec; new files use the new one. Partition evolution is metadata-only.
3. Which transform suits a high-cardinality user_id?
Show the answer
bucket(16, user_id). Bucketing hashes values into a fixed number of partitions.
Practice it
Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.
Go deeper
Primary sources: Iceberg: partitioning · Iceberg: partition evolution