Skip to content
Great engineers know 3 min read · All formats 1 practice problem ↓

Copy-on-write vs merge-on-read

Two ways to apply updates and deletes, and which one fits a given workload.

On this page

Show code in

Every code block on the page follows this.

You will learn

  • How copy-on-write applies updates and deletes
  • How merge-on-read applies them, and what readers must do
  • The read and write cost of each
  • How to configure each in Delta, Iceberg and Hudi

Read first

Comfortable with these? Read on.

TL;DR Copy-on-write rewrites every file that contains a changed row: writes are expensive, reads stay simple. Merge-on-read writes small files that record the changes and merges them when reading: writes are cheap, reads do more work until compaction folds the changes in.

Copy-on-write

To delete one row from a 512 MB file, copy-on-write reads the whole file, writes a new 512 MB file without that row, and commits "remove old file, add new file". Readers only ever see plain data files.

Merge-on-read

Merge-on-read leaves the data file alone and writes a small file saying "row 1,234 of file X is deleted" (or, for updates, a delete plus the new row version). Readers apply those deletes while scanning. Periodic compaction rewrites data files with the changes applied and drops the delete files.

Change · the update

Update the tier of customer 42, stored in a 512 MB file with 4 million rows.

The command

UPDATE customers SET tier = 'gold' WHERE customer_id = 42;

CoW · write

Read 512 MB, write a new 512 MB file with the updated row, commit. Cost: one full file rewrite per touched file.

MoR · write

Write a tiny delete file (or deletion vector) marking the old row, plus a small data file with the new row version. Cost: kilobytes.

MoR · read

Every reader of that file now also reads the delete file and filters the old row out. Thousands of updates mean thousands of small delete and data files until compaction runs.

Trade-offs

Copy-on-writeMerge-on-read
Write costHigh: rewrite touched filesLow: write small change files
Read costLowest: plain filesHigher: apply deletes while reading
Write latencySlow for scattered updatesFast, good for streaming upserts
MaintenanceCompaction for small files onlyRegular compaction needed to keep reads fast
Best forRead-heavy tables, batch updates, few changesFrequent updates and deletes, CDC, GDPR deletes

In each format

FormatHow
Delta LakeCopy-on-write by default. Deletion vectors (delta.enableDeletionVectors) give merge-on-read behaviour for DELETE, UPDATE and MERGE.
Apache IcebergPer operation: write.delete.mode, write.update.mode, write.merge.mode = copy-on-write or merge-on-read. MoR uses position or equality delete files, and deletion vectors in format v3.
Apache HudiChosen per table: COPY_ON_WRITE or MERGE_ON_READ table types. MoR stores base Parquet files plus row-based log files, merged by compaction.
Spark SQL
-- Iceberg: make updates and merges merge-on-read
ALTER TABLE db.customers SET TBLPROPERTIES (
  'write.update.mode' = 'merge-on-read',
  'write.merge.mode'  = 'merge-on-read',
  'write.delete.mode' = 'merge-on-read');

-- Delta: enable deletion vectors
ALTER TABLE customers SET TBLPROPERTIES ('delta.enableDeletionVectors' = true);
Why this matters: the choice is about where you pay: once at write time (CoW) or on every read until compaction (MoR). Tables updated constantly and read occasionally favour MoR; dashboards over rarely changing tables favour CoW.

Common mistakes

Using MoR without scheduled compaction

Reads slow down as delete files pile up.

Using CoW for streaming upserts

Each micro-batch rewrites many large files.

Forgetting reader compatibility

Older engines may not read delete files or deletion vectors.

Key takeaways

  • Copy-on-write rewrites touched files: costly writes, simple reads.
  • Merge-on-read writes small change files and merges at read time.
  • MoR needs regular compaction.
  • Delta uses deletion vectors; Iceberg sets modes per operation; Hudi per table type.

Check yourself

3 questions

1. Which is cheaper for frequent single-row updates?

Show the answer

Merge-on-read. MoR writes small delete/change files instead of rewriting whole files.

2. What must readers do with merge-on-read tables?

Show the answer

Apply delete files or deletion vectors while scanning. Changes are stored separately and merged at read time.

3. How does Delta Lake get merge-on-read behaviour?

Show the answer

Deletion vectors. Deletion vectors mark rows deleted without rewriting files.

Practice it

Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.

Solve: Apply a CDC Batch (MERGE Semantics) →

Go deeper

Primary sources: Iceberg: write properties · Hudi: table types · Delta: deletion vectors