Skip to content
Great engineers know 4 min read · Iceberg

Iceberg metadata tree

Metadata files, snapshots, manifest lists and manifests, and how query planning uses them.

You will learn

  • The layers of Iceberg metadata, from catalog to data file
  • How a commit creates a new snapshot
  • How query planning prunes manifests and files
  • How to inspect metadata with metadata tables

Read first

Comfortable with these? Read on.

TL;DR An Iceberg table is a tree: the catalog points to the current metadata file, which lists snapshots; each snapshot has a manifest list; each manifest list points to manifests, which list data files with per-column statistics. Commits add new nodes and atomically move the catalog pointer.

The metadata tree

Catalog:  db.events → s3://…/metadata/v43.metadata.json

v43.metadata.json          (JSON: schemas, partition specs, snapshot list, current snapshot)
└── snapshot 8274…  →  snap-8274-….avro          (manifest list)
    ├── manifest A.avro   partitions 2025-03-01..03-10, 1,200 files
    │   ├── data/….parquet  stats: rows, nulls, lower/upper bounds per column
    │   └── …
    └── manifest B.avro   partitions 2025-03-11..03-14, 300 files
        └── …
LayerFormatHolds
Catalog entryCatalog (REST, Glue, Hive, JDBC, Nessie...)The location of the current metadata file: the single source of truth
Metadata fileJSONSchemas (with column IDs), partition specs, sort orders, properties, the list of snapshots and which is current
Manifest listAvro, one per snapshotThe manifests in this snapshot, with partition value ranges and file counts per manifest
ManifestAvroData and delete files, each with partition values and column statistics
Data and delete filesParquet (or ORC, Avro)The rows, and in merge-on-read tables, deletions

How a commit works

Step 1 · write data

The writer writes new Parquet files. They are invisible: nothing references them.

Step 2 · new manifest

It writes a new manifest listing those files and their statistics. Unchanged manifests are reused, not rewritten.

Step 3 · new manifest list

It writes a manifest list for the new snapshot: the old manifests plus the new one.

Step 4 · new metadata file

It writes v44.metadata.json, which adds the new snapshot and marks it current.

Step 5 · swap

It asks the catalog to change the pointer from v43 to v44 only if it still points to v43. If another writer committed first, the swap fails, and the writer re-reads, checks for conflicts and retries on top of the new state.

Because each snapshot shares unchanged manifests with the previous one, commits write only a few small files regardless of table size.

How planning uses metadata

  1. 1
    Read the current metadata file from the catalog, then the snapshot's manifest list.
  2. 2
    Use the partition ranges stored per manifest to skip whole manifests that cannot match the filter.
  3. 3
    Read the remaining manifests and use per-file column bounds to skip data files.
  4. 4
    Hand the surviving files to the engine. No directory listing happens at any point.

That is why Iceberg scales to tables with millions of files: planning cost depends on the manifests touched, not on listing storage. It also means manifests need maintenance: many tiny manifests from frequent commits slow planning, which rewrite_manifests fixes.

Inspecting metadata

Spark SQL
-- Snapshot history
SELECT committed_at, snapshot_id, operation, summary FROM db.events.snapshots;

-- Data files with sizes and record counts
SELECT file_path, record_count, file_size_in_bytes FROM db.events.files;

-- Manifests and partitions
SELECT * FROM db.events.manifests;
SELECT * FROM db.events.partitions;
Note: Delta keeps a flat, ordered log of actions; Iceberg keeps a tree of immutable files. Both reach the same goal (an atomic, versioned list of files with statistics) with different trade-offs: Delta's log is simple and append-only; Iceberg's tree makes planning on huge tables cheap and relies on a catalog for the atomic step.

Common mistakes

Never maintaining manifests

Frequent commits create many small manifests and slow planning; rewrite them periodically.

Editing metadata files by hand

Breaks the immutable tree; use procedures.

Pointing two catalogs at one table

Each thinks it owns the pointer; commits can be lost.

Key takeaways

  • Catalog → metadata file → manifest list → manifests → data files.
  • A commit writes new metadata and atomically swaps the catalog pointer.
  • Planning prunes manifests by partition ranges and files by column bounds, with no listing.
  • Metadata tables (snapshots, files, manifests) make the tree queryable.

Check yourself

3 questions

1. What makes an Iceberg commit atomic?

Show the answer

The catalog's compare-and-swap of the current metadata pointer. Only the pointer swap must be atomic; everything else is new immutable files.

2. Which layer holds per-file column statistics?

Show the answer

Manifest. Manifests list data files with their bounds and counts.

3. Why does Iceberg avoid directory listing?

Show the answer

Manifests explicitly list every file in a snapshot. The file list comes from metadata, not storage listings.

Go deeper

Primary sources: Iceberg table spec · Iceberg: Spark queries and metadata tables