The small file problem
Why thousands of tiny files slow every read, how Spark jobs create them, and how to prevent and fix them.
On this page
Show code in
Every code block on the page follows this.
You will learn
- Why many small files slow down every read
- The five ways Spark jobs usually create them
- How to prevent them when writing
- How to fix them after the fact, and how to detect them
Why small files hurt
Reading a file has a fixed cost that does not depend on its size:
| Cost per file | Why it adds up |
|---|---|
| Listing | Object stores list about 1,000 keys per request. Listing 1 million files takes 1,000 sequential requests before any planning. |
| Open and footer read | Parquet keeps its schema and statistics at the end of each file; every file needs at least one extra request. |
| Metadata | For Hive tables, every partition and file is tracked by the metastore and the driver. For Delta and Iceberg, every file is an entry in the log or manifests, which grows and slows planning. |
| Tasks | Spark packs small files together (counting each as at least openCostInBytes, 4 MB) but scheduling still grows with the file count. |
| Compression and statistics | Tiny row groups compress worse, and min/max statistics over a few rows rarely let a reader skip anything useful. |
On object stores like S3, ADLS and GCS, request latency, not bandwidth, dominates. Reading 10 GB as 100,000 files can take ten times longer than the same 10 GB as 80 files.
How Spark jobs create them
- Too many shuffle partitions at write time. A write after a join or aggregation produces one file per non-empty partition. 200 partitions writing 50 MB of data gives 200 files of 250 KB.
- partitionBy on a high-cardinality column. Each task writes one file per partition value it holds. 200 tasks × 365 dates = up to 73,000 files per run.
- Streaming micro-batches. A stream that commits every 10 seconds writes files every 10 seconds: 8,640 batches a day, each with several files.
- Frequent small appends. Hourly or per-event loads of a few MB each.
- Over-partitioned sources. Reading thousands of files and writing straight back preserves the count.
Walk through one bad write
Step 1 · the job
event_date and country (60 countries). The upstream shuffle has 200 partitions.The command
daily.write.mode("append").partitionBy("event_date", "country").parquet(path)
Step 2 · the math
Step 3 · a year later
Step 4 · the fix
The command
daily.repartition("event_date", "country") \ .write.mode("append").partitionBy("event_date", "country").parquet(path)
INSERT INTO events_by_country SELECT /*+ REPARTITION(event_date, country) */ * FROM daily
Preventing them when writing
| Technique | How | Watch out |
|---|---|---|
| Let AQE coalesce | With AQE on, a write after a shuffle uses merged partitions near 64 MB | Only after a shuffle; raise advisoryPartitionSizeInBytes to 128-256 MB for write-heavy jobs |
| repartition(n) | Pick n from the output size: 50 GB / 256 MB ≈ 200 | A full shuffle |
| repartition(partition cols) | One task per table partition, so one file per directory | A huge partition value becomes one huge task (skew) |
| coalesce(n) | Merge without a shuffle | Reduces the parallelism of the whole stage |
| maxRecordsPerFile | .option("maxRecordsPerFile", 5_000_000) caps file size from above | Does not merge small files |
| Coarser partitioning | Partition by month instead of day, or not at all for small tables | See Partitioning done right |
Delta Lake can do this for you: optimized writes add an adaptive shuffle before writing so each partition gets fewer, larger files, and auto compaction runs a small compaction after a write. They are table properties (delta.autoOptimize.optimizeWrite, delta.autoOptimize.autoCompact) on Databricks and in recent open-source Delta releases; check your version.
Fixing them afterwards
# Plain Parquet: rewrite one partition into fewer files, then swap it in (spark.read.parquet(f"{path}/event_date=2025-03-01") .repartition(8) .write.mode("overwrite").parquet(f"{tmp}/event_date=2025-03-01")) # Delta Lake from delta.tables import DeltaTable DeltaTable.forPath(spark, path).optimize().executeCompaction()
-- Delta Lake OPTIMIZE events WHERE event_date >= '2025-03-01'; -- Apache Iceberg CALL catalog.system.rewrite_data_files(table => 'db.events');
With plain Parquet, rewriting in place is unsafe for concurrent readers: they can see the old files deleted before the new ones appear. Table formats make compaction a single atomic commit, which is one of their biggest practical benefits. Old files remain until VACUUM (Delta) or expire_snapshots (Iceberg) removes them.
Detecting the problem
DESCRIBE DETAIL tableon Delta returnsnumFilesandsizeInBytes: divide to get the average file size. Under about 32 MB is worth fixing.- For Iceberg, query the metadata table
db.events.filesforfile_size_in_bytes. - In the Spark UI, a scan with tens of thousands of tasks reading a few hundred KB each, or a long gap before the first job (planning and listing), points to small files.
Common mistakes
coalesce(1) to "fix" small files on big data
partitionBy on user_id or timestamp
Compacting plain Parquet in place while others read it
Key takeaways
- Each file has a fixed cost: listing, opening, metadata and often a task.
- Small files come from too many write partitions, high-cardinality partitionBy, streaming and small appends.
- Repartition by the partition columns, let AQE coalesce, and use maxRecordsPerFile to cap size.
- Table formats make compaction atomic; aim for 128 MB to 1 GB files.
Check yourself
3 questions1. 200 tasks each hold rows for 30 dates and write with partitionBy("date"). Up to how many files are written?
Show the answer
6,000. Each task writes one file per partition value it holds: 200 × 30.
2. Which change gives one file per date directory?
Show the answer
repartition("date") before the write. Repartitioning by the partition column sends all rows of a date to one task.
3. Why is compaction safer on Delta or Iceberg than on plain Parquet?
Show the answer
The rewrite is committed atomically, so readers see either the old files or the new ones. The new file list is published in one commit; old files stay readable until vacuumed.
Practice it
Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.
Go deeper
Compaction, OPTIMIZE and Z-orderLakehouse
Partitioning done rightSpark internals
Shuffle partitions
Primary sources: Performance tuning · Delta Lake: optimizations · Iceberg: maintenance