Deletion vectors
Mark rows as deleted without rewriting whole files, and the cost readers pay for it.
On this page
Show code in
Every code block on the page follows this.
You will learn
- What a deletion vector is and how it is stored
- Which operations use them, and the speed-up
- What readers must support
- How and when the deleted rows are physically removed
Read first
- Copy-on-write vs merge-on-read · 3 min read
- The Delta transaction log · 8 min read
Comfortable with these? Read on.
What it is
A deletion vector (DV) belongs to one data file and lists the positions of its deleted rows, encoded as a compressed RoaringBitmap. Small DVs can be stored inline in the commit; larger ones go into separate DV files. The add action for a file references its DV, so the log always says "this file, minus these rows".
Before · the table
part-07.parquet holds 3 million rows, 12 of which belong to a customer who asked to be deleted.Delete · the commit
The command
DELETE FROM events WHERE customer_id = 9001;
Read · the scan
Purge · later
OPTIMIZE rewrites files with DVs as part of compaction; REORG TABLE events APPLY (PURGE) rewrites them explicitly. After VACUUM, the old file is physically gone.Enabling and using
ALTER TABLE events SET TBLPROPERTIES ('delta.enableDeletionVectors' = true); -- Physically remove soft-deleted rows (e.g. for GDPR erasure), then vacuum REORG TABLE events APPLY (PURGE); VACUUM events;
- Supported for DELETE first (Delta 2.4), then UPDATE and MERGE in later Delta 3.x releases.
- Speed-ups are largest when changes touch few rows across many files: point deletes, GDPR requests, small CDC batches.
- Enabling DVs adds a table feature, raising the reader and writer protocol versions; readers must support deletion vectors to read the table at all.
In Iceberg
Iceberg format v2 uses position delete files (file path plus row position) and equality delete files (delete all rows matching these values). Format v3 adds binary deletion vectors stored in Puffin files, with at most one vector per data file, which is much cheaper to read than many position delete files. The two ecosystems converged on the same idea.
Costs
- Reads must load and apply DVs: cheap for a few, noticeable when most files carry large ones.
- Storage holds deleted rows until purge and vacuum, which matters for compliance deadlines.
- Compaction becomes part of the plan, not optional.
Common mistakes
Assuming DELETE erased the data
Enabling DVs while other engines read the table
Never compacting
Key takeaways
- A deletion vector is a bitmap of deleted row positions for one file.
- DELETE, UPDATE and MERGE write DVs instead of rewriting files.
- Readers must support DVs; enabling them upgrades the protocol.
- Purge with OPTIMIZE or REORG ... APPLY (PURGE), then VACUUM.
Check yourself
3 questions1. What does a deletion vector store?
Show the answer
Positions of deleted rows in one data file. It is a compact bitmap of row positions.
2. After a DELETE with DVs, are the rows physically gone?
Show the answer
No, until the files are rewritten and the old ones vacuumed. Rows stay in Parquet files until purge and vacuum.
3. What is Iceberg's v3 equivalent?
Show the answer
Deletion vectors stored in Puffin files. Format v3 introduces binary deletion vectors in Puffin files.
Go deeper
Primary sources: Delta: deletion vectors · Iceberg spec: deletion vectors