Schema enforcement and evolution
What a table rejects on write, how mergeSchema adds columns, and which changes are safe.
On this page
Show code in
Every code block on the page follows this.
Enforcement
A Delta or Iceberg write is checked against the table schema before commit. A plain ParquetParquet: A columnar file format: values of each column are stored together, so queries read only the columns they need. Learn more → folder accepts anything; a table format refuses:
| Write | Result |
|---|---|
| A column the table does not have | Rejected (unless schema merging is enabled) |
| A column with an incompatible type, e.g. string into int | Rejected |
| A missing nullable column | Accepted; the column is null for these rows |
| Column names differing only in case | Rejected in Delta (case-insensitive names must be unique) |
This is a feature: a producer that suddenly sends amount as text fails loudly at write time instead of corrupting months of data.
Adding columns
# Let this write add new columns (Delta) df.write.format("delta").mode("append").option("mergeSchema", "true").saveAsTable("events")
ALTER TABLE events ADD COLUMNS (device STRING, app_version STRING);
For MERGE, Delta supports automatic schema evolution with spark.databricks.delta.schema.autoMerge.enabled, and newer versions accept MERGE WITH SCHEMA EVOLUTION. Prefer explicit ALTER TABLE in production: schema changes become reviewed decisions instead of side effects of a write.
Which changes are safe
| Change | Delta Lake | Apache Iceberg |
|---|---|---|
| Add column | Yes | Yes |
| Widen type (int → long, float → double) | With type widening (preview in Delta 3.2, GA in 4.0) | Yes (int → long, float → double, decimal precision increase) |
| Rename column | Yes, with column mapping enabled | Yes |
| Drop column | Yes, with column mapping enabled | Yes |
| Reorder columns | Yes | Yes |
| Narrow or change type (long → int, string → int) | Rewrite the table (overwriteSchema) | Not allowed; rewrite |
Why renames are hard, and how IDs fix them
Parquet files identify columns by name. If a table renames email to contact_email by metadata only, old files still say email, and a name-based reader would see the new column as null. Iceberg solved this from the start by giving every column a permanent field ID stored in the Parquet files; readers match by ID, so names can change freely. Delta added the same idea as column mapping:
spark.sql(""" ALTER TABLE events SET TBLPROPERTIES ( 'delta.columnMapping.mode' = 'name', 'delta.minReaderVersion' = '2', 'delta.minWriterVersion' = '5') """) spark.sql("ALTER TABLE events RENAME COLUMN email TO contact_email") spark.sql("ALTER TABLE events DROP COLUMN legacy_flag")
ALTER TABLE events SET TBLPROPERTIES ( 'delta.columnMapping.mode' = 'name', 'delta.minReaderVersion' = '2', 'delta.minWriterVersion' = '5'); ALTER TABLE events RENAME COLUMN email TO contact_email; ALTER TABLE events DROP COLUMN legacy_flag;
Replacing the schema entirely
For incompatible changes, rewrite: df.write.mode("overwrite").option("overwriteSchema", "true").saveAsTable("events") in Delta, or CREATE OR REPLACE TABLE ... AS SELECT. This rewrites all data and keeps history, so time travel to the old schema still works within retention.
Common mistakes
- Turning on mergeSchema everywhere — Any upstream typo becomes a permanent new column.
- Renaming in Delta without column mapping — Not supported without it; enabling it changes the protocol.
- Assuming all engines read evolved tables — Protocol upgrades can lock out older readers.
What you learned
- What schema enforcement rejects, and why
- How to add columns with mergeSchema or ALTER TABLE
- Which column changes are safe, and how renames and drops work
- How Iceberg tracks columns by ID
Key takeaways
- Table formats reject writes that do not match the schema.
- Add columns with ALTER TABLE or mergeSchema; prefer explicit changes.
- Iceberg tracks columns by ID; Delta needs column mapping for renames and drops.
- Protocol upgrades for evolution features can break older readers.
Check yourself
3 questionsA Delta append contains a column the table lacks, without mergeSchema. What happens?
Show the answer
The write fails. Schema enforcement rejects unknown columns.
How does Iceberg allow renaming a column without rewriting files?
Show the answer
Columns are tracked by permanent field IDs, not names. Readers match columns by ID, so names in metadata can change.
What does Delta need to drop a column without rewriting data?
Show the answer
Column mapping. Column mapping decouples logical names from physical Parquet column names.
Practice it
Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.
Keep going
Up next · lesson 11 of 26 · 4 min readConstraints and generated columns
NOT NULL and CHECK constraints, generated columns and identity columns in Delta tables.
Related lessons
unionByNameData lake & lakehouse · 3 min read
Delta Lake 4.0Data lake & lakehouse · 4 min read
Iceberg metadata tree
Previous: MERGE INTO and upserts
Primary sources: Delta: schema validation and evolution · Delta: column mapping · Iceberg: schema evolution