Skip to content
Good engineers know 3 min read · Delta, Iceberg 1 practice problem ↓

Schema enforcement and evolution

What a table rejects on write, how mergeSchema adds columns, and which changes are safe.

TL;DR Table formats store the schema in metadata and reject writes that do not match it (enforcement). You can change the schema deliberately (evolution): add columns, widen types and, with column mapping in Delta or by default in Iceberg, rename and drop columns without rewriting data.
Read first: Open table formats (skip if you know it)

Enforcement

A Delta or Iceberg write is checked against the table schema before commit. A plain ParquetParquet: A columnar file format: values of each column are stored together, so queries read only the columns they need. Learn more → folder accepts anything; a table format refuses:

WriteResult
A column the table does not haveRejected (unless schema merging is enabled)
A column with an incompatible type, e.g. string into intRejected
A missing nullable columnAccepted; the column is null for these rows
Column names differing only in caseRejected in Delta (case-insensitive names must be unique)

This is a feature: a producer that suddenly sends amount as text fails loudly at write time instead of corrupting months of data.

Adding columns

PySparkSpark SQL · Two ways to add a column
# Let this write add new columns (Delta)
df.write.format("delta").mode("append").option("mergeSchema", "true").saveAsTable("events")
ALTER TABLE events ADD COLUMNS (device STRING, app_version STRING);

For MERGE, Delta supports automatic schema evolution with spark.databricks.delta.schema.autoMerge.enabled, and newer versions accept MERGE WITH SCHEMA EVOLUTION. Prefer explicit ALTER TABLE in production: schema changes become reviewed decisions instead of side effects of a write.

Which changes are safe

ChangeDelta LakeApache Iceberg
Add columnYesYes
Widen type (int → long, float → double)With type widening (preview in Delta 3.2, GA in 4.0)Yes (int → long, float → double, decimal precision increase)
Rename columnYes, with column mapping enabledYes
Drop columnYes, with column mapping enabledYes
Reorder columnsYesYes
Narrow or change type (long → int, string → int)Rewrite the table (overwriteSchema)Not allowed; rewrite

Why renames are hard, and how IDs fix them

Parquet files identify columns by name. If a table renames email to contact_email by metadata only, old files still say email, and a name-based reader would see the new column as null. Iceberg solved this from the start by giving every column a permanent field ID stored in the Parquet files; readers match by ID, so names can change freely. Delta added the same idea as column mapping:

PySparkSpark SQL
spark.sql("""
    ALTER TABLE events SET TBLPROPERTIES (
      'delta.columnMapping.mode' = 'name',
      'delta.minReaderVersion' = '2',
      'delta.minWriterVersion' = '5')
""")
spark.sql("ALTER TABLE events RENAME COLUMN email TO contact_email")
spark.sql("ALTER TABLE events DROP COLUMN legacy_flag")
ALTER TABLE events SET TBLPROPERTIES (
  'delta.columnMapping.mode' = 'name',
  'delta.minReaderVersion' = '2',
  'delta.minWriterVersion' = '5');
ALTER TABLE events RENAME COLUMN email TO contact_email;
ALTER TABLE events DROP COLUMN legacy_flag;
Watch out: enabling column mapping or type widening upgrades the table protocol. Older readers and some other engines may no longer be able to read the table. Check every consumer before changing protocol features.

Replacing the schema entirely

For incompatible changes, rewrite: df.write.mode("overwrite").option("overwriteSchema", "true").saveAsTable("events") in Delta, or CREATE OR REPLACE TABLE ... AS SELECT. This rewrites all data and keeps history, so time travel to the old schema still works within retention.

Common mistakes

  • Turning on mergeSchema everywhere — Any upstream typo becomes a permanent new column.
  • Renaming in Delta without column mapping — Not supported without it; enabling it changes the protocol.
  • Assuming all engines read evolved tables — Protocol upgrades can lock out older readers.

What you learned

  • What schema enforcement rejects, and why
  • How to add columns with mergeSchema or ALTER TABLE
  • Which column changes are safe, and how renames and drops work
  • How Iceberg tracks columns by ID

Key takeaways

  • Table formats reject writes that do not match the schema.
  • Add columns with ALTER TABLE or mergeSchema; prefer explicit changes.
  • Iceberg tracks columns by ID; Delta needs column mapping for renames and drops.
  • Protocol upgrades for evolution features can break older readers.

Check yourself

3 questions

A Delta append contains a column the table lacks, without mergeSchema. What happens?

Show the answer

The write fails. Schema enforcement rejects unknown columns.

How does Iceberg allow renaming a column without rewriting files?

Show the answer

Columns are tracked by permanent field IDs, not names. Readers match columns by ID, so names in metadata can change.

What does Delta need to drop a column without rewriting data?

Show the answer

Column mapping. Column mapping decouples logical names from physical Parquet column names.

Practice it

Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.

Solve: Combine Batches with Schema Drift →

Keep going

Up next · lesson 11 of 26 · 4 min read
Constraints and generated columns
NOT NULL and CHECK constraints, generated columns and identity columns in Delta tables.

Related lessons

Previous: MERGE INTO and upserts

Primary sources: Delta: schema validation and evolution · Delta: column mapping · Iceberg: schema evolution