Iceberg metadata tree
Metadata files, snapshots, manifest lists and manifests, and how query planning uses them.
On this page
Show code in
Every code block on the page follows this.
You will learn
- The layers of Iceberg metadata, from catalog to data file
- How a commit creates a new snapshot
- How query planning prunes manifests and files
- How to inspect metadata with metadata tables
Read first
- Open table formats · 3 min read
- ACID on object storage · 4 min read
Comfortable with these? Read on.
The metadata tree
Catalog: db.events → s3://…/metadata/v43.metadata.json
v43.metadata.json (JSON: schemas, partition specs, snapshot list, current snapshot)
└── snapshot 8274… → snap-8274-….avro (manifest list)
├── manifest A.avro partitions 2025-03-01..03-10, 1,200 files
│ ├── data/….parquet stats: rows, nulls, lower/upper bounds per column
│ └── …
└── manifest B.avro partitions 2025-03-11..03-14, 300 files
└── …| Layer | Format | Holds |
|---|---|---|
| Catalog entry | Catalog (REST, Glue, Hive, JDBC, Nessie...) | The location of the current metadata file: the single source of truth |
| Metadata file | JSON | Schemas (with column IDs), partition specs, sort orders, properties, the list of snapshots and which is current |
| Manifest list | Avro, one per snapshot | The manifests in this snapshot, with partition value ranges and file counts per manifest |
| Manifest | Avro | Data and delete files, each with partition values and column statistics |
| Data and delete files | Parquet (or ORC, Avro) | The rows, and in merge-on-read tables, deletions |
How a commit works
Step 1 · write data
Step 2 · new manifest
Step 3 · new manifest list
Step 4 · new metadata file
v44.metadata.json, which adds the new snapshot and marks it current.Step 5 · swap
Because each snapshot shares unchanged manifests with the previous one, commits write only a few small files regardless of table size.
How planning uses metadata
- 1Read the current metadata file from the catalog, then the snapshot's manifest list.
- 2Use the partition ranges stored per manifest to skip whole manifests that cannot match the filter.
- 3Read the remaining manifests and use per-file column bounds to skip data files.
- 4Hand the surviving files to the engine. No directory listing happens at any point.
That is why Iceberg scales to tables with millions of files: planning cost depends on the manifests touched, not on listing storage. It also means manifests need maintenance: many tiny manifests from frequent commits slow planning, which rewrite_manifests fixes.
Inspecting metadata
-- Snapshot history SELECT committed_at, snapshot_id, operation, summary FROM db.events.snapshots; -- Data files with sizes and record counts SELECT file_path, record_count, file_size_in_bytes FROM db.events.files; -- Manifests and partitions SELECT * FROM db.events.manifests; SELECT * FROM db.events.partitions;
Common mistakes
Never maintaining manifests
Editing metadata files by hand
Pointing two catalogs at one table
Key takeaways
- Catalog → metadata file → manifest list → manifests → data files.
- A commit writes new metadata and atomically swaps the catalog pointer.
- Planning prunes manifests by partition ranges and files by column bounds, with no listing.
- Metadata tables (snapshots, files, manifests) make the tree queryable.
Check yourself
3 questions1. What makes an Iceberg commit atomic?
Show the answer
The catalog's compare-and-swap of the current metadata pointer. Only the pointer swap must be atomic; everything else is new immutable files.
2. Which layer holds per-file column statistics?
Show the answer
Manifest. Manifests list data files with their bounds and counts.
3. Why does Iceberg avoid directory listing?
Show the answer
Manifests explicitly list every file in a snapshot. The file list comes from metadata, not storage listings.
Go deeper
Hidden partitioning and partition evolutionLakehouse
The Delta transaction logLakehouse
Catalogs: Unity, Polaris, Iceberg REST
Primary sources: Iceberg table spec · Iceberg: Spark queries and metadata tables