Skip to content
Everyone knows 3 min read · All formats 1 practice problem ↓

Open table formats

Delta Lake, Iceberg and Hudi: a metadata layer that turns a folder of Parquet files into a table.

You will learn

  • What a table format adds on top of Parquet files
  • Where Delta Lake, Iceberg and Hudi came from, and how they differ
  • What all three have in common
  • How to choose between them

Read first

Comfortable with these? Read on.

TL;DR A table format is a specification for metadata that lists which data files make up each version of a table, plus its schema and partitioning. Delta Lake, Apache Iceberg and Apache Hudi all store data as Parquet and differ mostly in how their metadata is organised and committed.

What a table format is

Without a table format, a table is "whatever files are in this folder". With one, a table is "whatever files the current metadata lists". That change makes commits atomic (switch to new metadata in one step), allows old versions to stay readable, and lets engines plan queries from metadata instead of listing storage.

The three formats

Delta LakeApache IcebergApache Hudi
OriginDatabricks, open-sourced 2019, now a Linux Foundation projectNetflix, donated to ApacheUber, donated to Apache
Original focusReliable Spark pipelines and streamingHuge tables, engine independence, correct schema and partition evolutionIncremental upserts and fast ingestion
MetadataOrdered JSON commit log plus Parquet checkpointsTree: metadata file, manifest lists, manifestsTimeline of commits in a .hoodie folder
CommitCreate the next numbered log file atomicallyAtomically swap the catalog's pointer to a new metadata fileComplete an instant on the timeline
Strong pointsTight Spark integration, simple model, wide adoption on DatabricksBroad engine support (Spark, Trino, Flink, Snowflake, BigQuery, Athena...), hidden partitioningRecord-level indexes, merge-on-read tables, built-in table services

What they share

  • Parquet data files (Hudi also uses Avro log files for merge-on-read).
  • ACID transactions with optimistic concurrency.
  • Schema enforcement and evolution.
  • Time travel to earlier versions or snapshots.
  • Updates, deletes and MERGE, by copy-on-write or merge-on-read.
  • File-level statistics for data skipping.
  • Maintenance operations: compaction and removal of old files.

The formats have converged a lot. Features once unique to one (deletion vectors, row-level changes, clustering) now exist in several, and tools like Delta UniForm and Apache XTable let one copy of data be read as more than one format.

How teams choose

If you...Often a good fit
Run mostly on Databricks and SparkDelta Lake
Need many engines and vendors on the same tablesIceberg (the broadest support today)
Ingest high-volume CDC upserts with low latencyHudi, or Delta / Iceberg with merge-on-read
Use a specific platformIts native format: Snowflake, AWS and Google favour Iceberg; Databricks and Microsoft Fabric favour Delta
Note: in interviews, there is rarely a single right answer. Show you know the trade-offs: metadata model, engine support, write patterns, and the platform your company already runs.

Common mistakes

Treating the formats as file formats

All three use Parquet files; the difference is the metadata.

Writing to a table with an engine that ignores the format

Plain Parquet writes bypass the log and corrupt the table view.

Choosing on benchmarks alone

Engine support and operations matter more for most teams.

Key takeaways

  • A table format is metadata that defines which files form each table version.
  • Delta uses a commit log, Iceberg a metadata tree, Hudi a timeline.
  • All three offer ACID, time travel, schema evolution and row-level changes.
  • Choose by engines, write patterns and platform, not benchmarks alone.

Check yourself

3 questions

1. What file format do all three table formats use for data?

Show the answer

Parquet. All three store data in Parquet (Hudi also uses Avro logs for merge-on-read).

2. How does Iceberg make a commit atomic?

Show the answer

The catalog atomically swaps the pointer to a new metadata file. Iceberg relies on an atomic compare-and-swap in the catalog.

3. Which format originated at Uber for incremental upserts?

Show the answer

Hudi. Hudi (Hadoop Upserts Deletes and Incrementals) came from Uber.

Practice it

Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.

Solve: Rebuild a Table from Its Transaction Log →

Go deeper

Primary sources: Delta Lake · Apache Iceberg · Apache Hudi