Skip to content
Great engineers know 5 min read · Iceberg, Delta

Branches, tags and write-audit-publish

A bad batch that lands in a production table is read by every dashboard within minutes. Write-audit-publish stages a write somewhere invisible, checks it, and only then makes it visible. Iceberg branches make that a first-class feature.

Read first: Iceberg metadata tree · Time travel and RESTORE (skip if you know them)

Write, audit, publish

  1. Write

    Commit the new data where production readers cannot see it

  2. Audit

    Run checks on it: row counts, nulls, duplicates, reconciliation with the source

  3. Publish

    Make it visible atomically, or discard it if checks fail

Without WAP, quality checks run after the data is already live, and a failed check means a rollback after people may have read bad numbers.

Iceberg branches and tags

Every Iceberg table already has a history of snapshotssnapshot: A complete, consistent version of a table at one moment. Readers always see one snapshot, never a half-finished write. Learn more →. A branch is a named, movable pointer to a snapshot that can receive its own commits; main is the branch normal reads and writes use. A tag is a named, fixed pointer to one snapshot, useful for "the data as of quarter close". Both can carry retention settings, and snapshots they reference are kept when old snapshots expire.

PySparkSpark SQL · Branches and tags (Spark SQL with the Iceberg extensions)
# a tag: keep the quarter-end state for a year
spark.sql("ALTER TABLE prod.db.sales CREATE TAG `q4_2024` RETAIN 365 DAYS")
spark.sql("SELECT * FROM prod.db.sales VERSION AS OF 'q4_2024'")

# a branch: an isolated line of commits
spark.sql("ALTER TABLE prod.db.sales CREATE BRANCH audit")
spark.sql("SELECT * FROM prod.db.sales VERSION AS OF 'audit'")
# or: SELECT * FROM prod.db.sales.branch_audit;
-- a tag: keep the quarter-end state for a year
ALTER TABLE prod.db.sales CREATE TAG `q4_2024` RETAIN 365 DAYS;
SELECT * FROM prod.db.sales VERSION AS OF 'q4_2024';

-- a branch: an isolated line of commits
ALTER TABLE prod.db.sales CREATE BRANCH audit;
SELECT * FROM prod.db.sales VERSION AS OF 'audit';
-- or: SELECT * FROM prod.db.sales.branch_audit;

WAP with a branch, step by step

Step 1 · branch

Enable WAP on the table once, then create a branch from the current state of main, with a retention so a forgotten branch is cleaned up.

The command

spark.sql("ALTER TABLE prod.db.sales SET TBLPROPERTIES ('write.wap.enabled' = 'true')")
spark.sql("ALTER TABLE prod.db.sales CREATE BRANCH etl_20250301 RETAIN 7 DAYS")
ALTER TABLE prod.db.sales SET TBLPROPERTIES ('write.wap.enabled' = 'true');
ALTER TABLE prod.db.sales CREATE BRANCH etl_20250301 RETAIN 7 DAYS;

Step 2 · write

Point the session at the branch with spark.wap.branch and run the job unchanged. Its commits go to the branch; readers of main see nothing.

The command

spark.conf.set("spark.wap.branch", "etl_20250301")
daily.writeTo("prod.db.sales").append()

Step 3 · audit

Query the branch and run the checks. Compare with main to see exactly what the job would change.

The command

spark.sql("""
    SELECT COUNT(*), COUNT_IF(amount IS NULL)
    FROM prod.db.sales VERSION AS OF 'etl_20250301'
    WHERE sale_date = DATE'2025-03-01'
""")
SELECT COUNT(*), COUNT_IF(amount IS NULL)
FROM prod.db.sales VERSION AS OF 'etl_20250301'
WHERE sale_date = DATE'2025-03-01';

Step 4 · publish

If the checks pass, fast-forward main to the branch: one atomic metadata change. If they fail, drop the branch; main was never touched.

The command

spark.sql("CALL prod.system.fast_forward('db.sales', 'main', 'etl_20250301')")
spark.sql("ALTER TABLE prod.db.sales DROP BRANCH etl_20250301")
CALL prod.system.fast_forward('db.sales', 'main', 'etl_20250301');
ALTER TABLE prod.db.sales DROP BRANCH etl_20250301;

Fast-forward only works if main has not moved since the branch was created; otherwise publish by cherry-picking the branch's snapshot, or redo the write. Before branches existed, Iceberg did WAP with write.wap.enabled and a spark.wap.id on staged snapshots, published with the cherrypick_snapshot procedure; you may still see it.

On Delta Lake

Delta has no table branches. The same guarantee comes from staging:

  • Staging table: write the batch to a staging table, audit it, then publish with one MERGE or INSERT OVERWRITE ... REPLACE WHERE into production. The publish step is a single atomic commit.
  • Shallow clone: CREATE TABLE sales_wap SHALLOW CLONE sales copies only metadata. Run the job against the clone, audit, then apply the same change to production. Useful for testing a whole pipeline against production data without copying it.
  • Rollback as a safety net: if bad data does get published, RESTORE to the previous version, but readers may already have seen it, which is what WAP avoids.

Branching the whole catalog

Iceberg branches are per table. A pipeline that updates five tables needs all five to publish together. Catalog-level tools give Git-like branches across many tables: Project Nessie (an Iceberg catalog with branches, tags and multi-table commits) and lakeFS (branches over object storage for any format). The pipeline writes all tables on a branch, audits, and merges the branch: every table changes at once.

Common mistakes

  • Running quality checks after publishing — Readers see bad data before the check fails.
  • Forgetting to drop failed branches — Their snapshots are retained and storage grows.
  • Assuming fast-forward always works — If main moved, fast-forward fails; cherry-pick or rewrite.
  • Publishing related tables one by one — Readers see a mix of old and new across tables. Use catalog-level branches or publish in one transaction where supported.

Interview prep

THE QUESTION

"How do you make sure a bad batch never becomes visible in a production table?"

Write-audit-publish. On Iceberg: enable write.wap.enabled, create a branch, set spark.wap.branch so the job commits to it, run the audits against the branch, then fast-forward main to it (atomic), or drop the branch if checks fail. On Delta: write to a staging table or shallow clone, audit, then publish with one atomic MERGE or REPLACE WHERE. For several tables that must change together, use catalog-level branching (Nessie, lakeFS). Mention tags for keeping audited snapshots, and RESTORE only as a fallback because readers may already have seen the data.

Avoid saying: "run the checks after the write and RESTORE if they fail". Readers may already have used the bad data.

What the interviewer asks next. Answer out loud first, then open the strong answer.

Follow-up"Why use a tag instead of just remembering a snapshot id?"
A tag is a named reference that is kept when old snapshots expire (with its own retention), so VERSION AS OF 'q4_2024' keeps working. A remembered snapshot id stops working once expire_snapshots removes it.
Scenario"Another job committed to main while your audit was running. What happens at publish?"
fast_forward fails because main is no longer an ancestor of the branch head. Either cherry-pick the branch's snapshot onto main (if it is an append and does not conflict), or rerun the write on a fresh branch from the new main.
Trap"Iceberg branches give you multi-table transactions."
Branches are per table. Publishing five branches is five separate commits. For atomic multi-table changes, use a catalog with multi-table commits such as Nessie, or lakeFS.

What you learned

  • What write-audit-publish (WAP) is and why it matters
  • How Iceberg branches and tags work
  • How to run WAP with an Iceberg branch, step by step
  • How to get the same guarantee on Delta, and with catalog-level branching

Key takeaways

  • WAP stages a write, checks it, and only then makes it visible.
  • Iceberg branches take commits in isolation; tags pin a snapshot.
  • With write.wap.enabled and spark.wap.branch, a job writes to a branch unchanged; fast_forward publishes atomically.
  • On Delta, use a staging table or shallow clone and one atomic publish; Nessie or lakeFS branch many tables at once.

Check yourself

3 questions

What is the difference between an Iceberg branch and a tag?

Show the answer

A branch can receive new commits; a tag points at one fixed snapshot. Branches move with commits; tags are fixed.

How does a Spark job write to an Iceberg branch without code changes?

Show the answer

Setting spark.wap.branch in the session. Writes in the session then commit to that branch.

How do you publish a WAP branch to main?

Show the answer

fast_forward main to the branch (or cherry-pick its snapshot). It is an atomic metadata update of the main pointer.

Keep going

Up next · lesson 27 of 30 · 4 min read
Catalogs: Unity, Polaris, Iceberg REST
What a catalog does for a lakehouse, and the open catalogs that arrived in 2024 and 2025.

Related lessons

Previous: Concurrency and conflicts

Primary sources: Iceberg branching and tagging · Iceberg Spark procedures · Iceberg Spark DDL: branches and tags