Open table formats
Delta Lake, Iceberg and Hudi: a metadata layer that turns a folder of Parquet files into a table.
On this page
You will learn
- What a table format adds on top of Parquet files
- Where Delta Lake, Iceberg and Hudi came from, and how they differ
- What all three have in common
- How to choose between them
Read first
- Lake, warehouse, lakehouse · 3 min read
- Parquet and columnar storage · 4 min read
Comfortable with these? Read on.
What a table format is
Without a table format, a table is "whatever files are in this folder". With one, a table is "whatever files the current metadata lists". That change makes commits atomic (switch to new metadata in one step), allows old versions to stay readable, and lets engines plan queries from metadata instead of listing storage.
The three formats
| Delta Lake | Apache Iceberg | Apache Hudi | |
|---|---|---|---|
| Origin | Databricks, open-sourced 2019, now a Linux Foundation project | Netflix, donated to Apache | Uber, donated to Apache |
| Original focus | Reliable Spark pipelines and streaming | Huge tables, engine independence, correct schema and partition evolution | Incremental upserts and fast ingestion |
| Metadata | Ordered JSON commit log plus Parquet checkpoints | Tree: metadata file, manifest lists, manifests | Timeline of commits in a .hoodie folder |
| Commit | Create the next numbered log file atomically | Atomically swap the catalog's pointer to a new metadata file | Complete an instant on the timeline |
| Strong points | Tight Spark integration, simple model, wide adoption on Databricks | Broad engine support (Spark, Trino, Flink, Snowflake, BigQuery, Athena...), hidden partitioning | Record-level indexes, merge-on-read tables, built-in table services |
What they share
- Parquet data files (Hudi also uses Avro log files for merge-on-read).
- ACID transactions with optimistic concurrency.
- Schema enforcement and evolution.
- Time travel to earlier versions or snapshots.
- Updates, deletes and MERGE, by copy-on-write or merge-on-read.
- File-level statistics for data skipping.
- Maintenance operations: compaction and removal of old files.
The formats have converged a lot. Features once unique to one (deletion vectors, row-level changes, clustering) now exist in several, and tools like Delta UniForm and Apache XTable let one copy of data be read as more than one format.
How teams choose
| If you... | Often a good fit |
|---|---|
| Run mostly on Databricks and Spark | Delta Lake |
| Need many engines and vendors on the same tables | Iceberg (the broadest support today) |
| Ingest high-volume CDC upserts with low latency | Hudi, or Delta / Iceberg with merge-on-read |
| Use a specific platform | Its native format: Snowflake, AWS and Google favour Iceberg; Databricks and Microsoft Fabric favour Delta |
Common mistakes
Treating the formats as file formats
Writing to a table with an engine that ignores the format
Choosing on benchmarks alone
Key takeaways
- A table format is metadata that defines which files form each table version.
- Delta uses a commit log, Iceberg a metadata tree, Hudi a timeline.
- All three offer ACID, time travel, schema evolution and row-level changes.
- Choose by engines, write patterns and platform, not benchmarks alone.
Check yourself
3 questions1. What file format do all three table formats use for data?
Show the answer
Parquet. All three store data in Parquet (Hudi also uses Avro logs for merge-on-read).
2. How does Iceberg make a commit atomic?
Show the answer
The catalog atomically swaps the pointer to a new metadata file. Iceberg relies on an atomic compare-and-swap in the catalog.
3. Which format originated at Uber for incremental upserts?
Show the answer
Hudi. Hudi (Hadoop Upserts Deletes and Incrementals) came from Uber.
Practice it
Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.
Go deeper
Primary sources: Delta Lake · Apache Iceberg · Apache Hudi