Iceberg format v3
Deletion vectors, row lineage, default values and VARIANT: what the v3 table spec adds.
On this page
Show code in
Every code block on the page follows this.
You will learn
- What a table format version is, and why v3 matters
- The main additions: deletion vectors, row lineage, default values, new types
- What readers and writers must support
- How to adopt v3 safely
Read first
- Iceberg metadata tree · 4 min read
- Deletion vectors · 3 min read
Comfortable with these? Read on.
Format versions
The Iceberg spec is versioned. v1 defined analytic tables with snapshots and manifests. v2 added row-level deletes (position and equality delete files) for merge-on-read. v3 is the next step; a table declares its format-version, and engines refuse to write a version they do not understand.
What v3 adds
| Feature | What it does | Why it matters |
|---|---|---|
| Deletion vectors | Binary bitmaps of deleted rows stored in Puffin files, at most one per data file | Cheaper merge-on-read than many position delete files; aligns with Delta's design |
| Row lineage | Each row gets a stable row id and the sequence number of its last update | Efficient change tracking and incremental processing without diffing snapshots |
| Default values | Columns can declare an initial default (for existing rows) and a write default | Add a NOT NULL column with a default without rewriting data |
| VARIANT type | Semi-structured data in a binary encoding, shared with Spark and Parquet | Store JSON-like data efficiently (see The VARIANT type) |
| Geospatial types | geometry and geography types | Spatial data as first-class columns |
| Nanosecond timestamps | timestamp_ns and timestamptz_ns | Precise event times from systems that produce nanoseconds |
| Multi-argument transforms | Partition and sort transforms over several columns | More flexible layouts |
Row lineage in practice
With row lineage, an engine can ask "which rows changed since sequence number 1,042?" and get exact answers from metadata columns, instead of comparing two snapshots row by row. That makes incremental pipelines and change feeds over Iceberg much cheaper, similar to what Delta's Change Data Feed offers.
Adopting v3
-- New table CREATE TABLE db.events (...) USING iceberg TBLPROPERTIES ('format-version' = '3'); -- Upgrade an existing table (one way: cannot downgrade) ALTER TABLE db.events SET TBLPROPERTIES ('format-version' = '3');
- Check every engine that reads or writes the table: Spark with a recent Iceberg runtime, Trino, Flink, warehouses. A reader without v3 support cannot read the table after the upgrade.
- Upgrading is one-way.
- Existing v2 delete files remain valid; new deletes can use deletion vectors.
- Support arrives feature by feature in engine releases; check release notes for which v3 features each engine implements.
Common mistakes
Upgrading before all readers support v3
Assuming every engine supports every v3 feature
Confusing table format version with library version
Key takeaways
- Iceberg v3 is a new table spec version, set per table with format-version.
- Adds deletion vectors, row lineage, default values and new types including VARIANT.
- The upgrade is one-way; every engine must support v3.
- Many v3 features converge with Delta Lake.
Check yourself
3 questions1. Where are v3 deletion vectors stored?
Show the answer
In Puffin files. v3 stores binary deletion vectors in Puffin files.
2. What does row lineage provide?
Show the answer
Stable row ids and the sequence number of each row's last update. It enables exact change tracking from metadata.
3. Can a table be downgraded from v3 to v2?
Show the answer
No, the upgrade is one-way. Format version upgrades are not reversible.
Go deeper
Primary sources: Iceberg table spec