Copy-on-write vs merge-on-read
Two ways to apply updates and deletes, and which one fits a given workload.
On this page
Show code in
Every code block on the page follows this.
You will learn
- How copy-on-write applies updates and deletes
- How merge-on-read applies them, and what readers must do
- The read and write cost of each
- How to configure each in Delta, Iceberg and Hudi
Copy-on-write
To delete one row from a 512 MB file, copy-on-write reads the whole file, writes a new 512 MB file without that row, and commits "remove old file, add new file". Readers only ever see plain data files.
Merge-on-read
Merge-on-read leaves the data file alone and writes a small file saying "row 1,234 of file X is deleted" (or, for updates, a delete plus the new row version). Readers apply those deletes while scanning. Periodic compaction rewrites data files with the changes applied and drops the delete files.
Change · the update
The command
UPDATE customers SET tier = 'gold' WHERE customer_id = 42;
CoW · write
MoR · write
MoR · read
Trade-offs
| Copy-on-write | Merge-on-read | |
|---|---|---|
| Write cost | High: rewrite touched files | Low: write small change files |
| Read cost | Lowest: plain files | Higher: apply deletes while reading |
| Write latency | Slow for scattered updates | Fast, good for streaming upserts |
| Maintenance | Compaction for small files only | Regular compaction needed to keep reads fast |
| Best for | Read-heavy tables, batch updates, few changes | Frequent updates and deletes, CDC, GDPR deletes |
In each format
| Format | How |
|---|---|
| Delta Lake | Copy-on-write by default. Deletion vectors (delta.enableDeletionVectors) give merge-on-read behaviour for DELETE, UPDATE and MERGE. |
| Apache Iceberg | Per operation: write.delete.mode, write.update.mode, write.merge.mode = copy-on-write or merge-on-read. MoR uses position or equality delete files, and deletion vectors in format v3. |
| Apache Hudi | Chosen per table: COPY_ON_WRITE or MERGE_ON_READ table types. MoR stores base Parquet files plus row-based log files, merged by compaction. |
-- Iceberg: make updates and merges merge-on-read ALTER TABLE db.customers SET TBLPROPERTIES ( 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read', 'write.delete.mode' = 'merge-on-read'); -- Delta: enable deletion vectors ALTER TABLE customers SET TBLPROPERTIES ('delta.enableDeletionVectors' = true);
Common mistakes
Using MoR without scheduled compaction
Using CoW for streaming upserts
Forgetting reader compatibility
Key takeaways
- Copy-on-write rewrites touched files: costly writes, simple reads.
- Merge-on-read writes small change files and merges at read time.
- MoR needs regular compaction.
- Delta uses deletion vectors; Iceberg sets modes per operation; Hudi per table type.
Check yourself
3 questions1. Which is cheaper for frequent single-row updates?
Show the answer
Merge-on-read. MoR writes small delete/change files instead of rewriting whole files.
2. What must readers do with merge-on-read tables?
Show the answer
Apply delete files or deletion vectors while scanning. Changes are stored separately and merged at read time.
3. How does Delta Lake get merge-on-read behaviour?
Show the answer
Deletion vectors. Deletion vectors mark rows deleted without rewriting files.
Practice it
Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.
Go deeper
Primary sources: Iceberg: write properties · Hudi: table types · Delta: deletion vectors