Lakehouse
23 in-depth lessons. Delta Lake, Apache Iceberg and Apache Hudi: how open table formats bring transactions to object storage, how to run them, and what is new.
Everyone knows
- Lake, warehouse, lakehouse · Everyone knows · 3 min read. What data lakes and warehouses each get right, and what the lakehouse borrows from both.
- Parquet and columnar storage · Everyone knows · 4 min read. Row groups, column chunks and statistics: why columnar files let engines skip most of the data.
- Open table formats · Everyone knows · 3 min read. Delta Lake, Iceberg and Hudi: a metadata layer that turns a folder of Parquet files into a table.
- ACID on object storage · Everyone knows · 4 min read. Why plain Parquet folders break under concurrent writes and failures, and how table formats fix it.
- The medallion architecture · Everyone knows · 3 min read. Bronze, silver and gold layers: what belongs in each, and where the pattern goes wrong.
Good engineers know
- The Delta transaction log · Good engineers know · 8 min read. Commit files, checkpoints and optimistic concurrency: how the _delta_log turns files into a table.
- Time travel and RESTORE · Good engineers know · 3 min read. Query or roll back to any earlier version of a table, and what retention does to that promise.
- MERGE INTO and upserts · Good engineers know · 4 min read. Matched, not-matched and not-matched-by-source clauses, step by step, and the duplicate-source error.
- Schema enforcement and evolution · Good engineers know · 3 min read. What a table rejects on write, how mergeSchema adds columns, and which changes are safe.
- Compaction, OPTIMIZE and Z-order · Good engineers know · 3 min read. Rewrite small files into large ones and cluster data so queries skip more of it.
- VACUUM and retention · Good engineers know · 4 min read. Delete old data files safely without breaking running readers or time travel.
- Change Data Feed · Good engineers know · 3 min read. Read the row-level inserts, updates and deletes made to a Delta table between two versions.
- Partitioning done right · Good engineers know · 4 min read. When to partition a table, how to choose the column, and how over-partitioning backfires.
Great engineers know
- Iceberg metadata tree · Great engineers know · 4 min read. Metadata files, snapshots, manifest lists and manifests, and how query planning uses them.
- Hidden partitioning and partition evolution · Great engineers know · 3 min read. Iceberg partition transforms, and changing a table layout without rewriting its data.
- Copy-on-write vs merge-on-read · Great engineers know · 3 min read. Two ways to apply updates and deletes, and which one fits a given workload.
- Deletion vectors · Great engineers know · 3 min read. Mark rows as deleted without rewriting whole files, and the cost readers pay for it.
- Liquid clustering · Great engineers know · 3 min read. Incremental clustering in Delta Lake that replaces fixed partitions and Z-order.
- Concurrency and conflicts · Great engineers know · 3 min read. Optimistic concurrency, the conflicts that fail a commit, and how to design writers around them.
- Catalogs: Unity, Polaris, Iceberg REST · Great engineers know · 3 min read. What a catalog does for a lakehouse, and the open catalogs that arrived in 2024 and 2025.
- Format interoperability · Great engineers know · 3 min read. Delta UniForm and Apache XTable: one copy of the data, readable as Delta, Iceberg or Hudi.
- Iceberg format v3 · Great engineers know · 3 min read. Deletion vectors, row lineage, default values and VARIANT: what the v3 table spec adds.
- Delta Lake 4.0 · Great engineers know · 3 min read. VARIANT, type widening, collations and the other changes that shipped alongside Spark 4.0.