Skip to content
Everyone knows 4 min read · All formats 1 practice problem ↓

ACID on object storage

Why plain Parquet folders break under concurrent writes and failures, and how table formats fix it.

You will learn

  • What each letter of ACID means for a table on S3
  • Why plain Parquet folders fail each property
  • How table formats provide atomic commits on object storage
  • What object storage itself guarantees, and what it does not

Read first

Comfortable with these? Read on.

TL;DR Object stores give durable, strongly consistent single-object writes, but no multi-file transactions and no atomic rename of folders. Table formats get atomicity by making a commit a single small write (a log file or a metadata pointer), so a set of data files becomes visible all at once or not at all.

ACID, applied to tables

PropertyMeansPlain Parquet folder
AtomicityA write is all or nothingA job failing after writing 60 of 100 files leaves 60 visible files
ConsistencyData always meets the table's rules (schema, constraints)Any job can write any schema into the folder
IsolationConcurrent readers and writers do not see each other's partial workA reader listing during a write sees some new files and some old
DurabilityCommitted data survives failuresProvided by the object store itself

What object storage guarantees

  • Durable: objects are replicated; S3 is designed for 99.999999999% durability.
  • Strongly consistent: since December 2020, S3 gives read-after-write consistency for all operations, including listings. ADLS and GCS are also strongly consistent.
  • Atomic per object: a single PUT either fully succeeds or does not happen.
  • No atomic rename: S3 has no real directories; "renaming" a folder copies and deletes every object, which is slow and not atomic. HDFS-era commit protocols that relied on renaming a temporary folder are slow and unsafe on S3.
  • No multi-object transactions: nothing lets you publish 100 files atomically.

The trick: one small atomic write

Table formats never need to make 100 files appear atomically. They write the data files first, invisible because no metadata references them, then perform one atomic operation that switches the table to a new file list.

Write · data files

The job writes 100 new Parquet files with unique names. Readers ignore them: the current metadata does not list them.

Commit · one atomic step

Delta: create _delta_log/00000000000000000043.json only if it does not already exist. Iceberg: ask the catalog to swap the table pointer from metadata v42 to v43 only if it still points to v42.

Success · visible at once

Readers that start now see version 43 with all 100 files. Readers already running continue on version 42, unaffected: that is snapshot isolation.

Failure · nothing visible

If the job dies before the commit, the 100 files are orphans that no version references. Readers never see them; a cleanup job deletes them later.

The commit step needs a "create if absent" or compare-and-swap primitive. HDFS and ADLS offer atomic rename; GCS offers conditional writes; S3 added conditional writes (If-None-Match) in 2024. Before that, multi-writer Delta tables on S3 needed an external coordinator such as DynamoDB, and Iceberg relies on its catalog (Glue, a REST catalog, a database) for the atomic swap.

Consistency and isolation in practice

  • Schema enforcement rejects writes that do not match the table schema, and constraints such as NOT NULL or CHECK (Delta) are validated before commit.
  • Readers get snapshot isolation: each query reads one consistent version.
  • Writers use optimistic concurrency: they check at commit time whether a concurrent commit conflicts with theirs, and retry or fail. See Concurrency and conflicts.

Common mistakes

Relying on folder renames on S3

Not atomic and slow; use a table format or a proper committer.

Deleting data files directly in storage

The metadata still lists them; reads fail.

Assuming multi-writer safety everywhere

Check that your storage and catalog provide the needed atomic primitive.

Key takeaways

  • Object stores are durable and consistent per object but have no multi-file transactions or atomic renames.
  • Table formats write data files first and publish them with one atomic metadata write.
  • Readers get snapshot isolation; failed jobs leave only invisible orphan files.
  • The atomic step needs put-if-absent or a catalog compare-and-swap.

Check yourself

3 questions

1. A job crashes after writing data files but before committing. What do readers of a Delta or Iceberg table see?

Show the answer

The previous version, unchanged. Uncommitted files are not referenced by any version.

2. Why are rename-based commits a problem on S3?

Show the answer

A folder rename is a copy and delete of every object: slow and not atomic. S3 has no real directories, so renaming a folder is many separate operations.

3. Which isolation do readers get?

Show the answer

Snapshot isolation: each query sees one committed version. Queries read a fixed version of the metadata.

Practice it

Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.

Solve: Rebuild a Table from Its Transaction Log →

Go deeper

Primary sources: Delta Lake: storage configuration · Amazon S3 consistency