Caching and persistence
cache(), persist() and storage levels: when caching speeds a job up and when it hurts.
On this page
Show code in
Every code block on the page follows this.
You will learn
- What cache() and persist() actually do, and when
- The storage levels and their trade-offs
- When caching speeds a job up and when it slows it down
- How to release cached data
cache() marks a DataFrame to be kept in executor memory (spilling to disk) the first time an action computes it, so later actions reuse it. It helps only when the same data is used several times, and it costs memory that joins and aggregations also need.What it does
Without caching, every action recomputes a DataFrame from its source. df.cache() tells Spark: the next time you compute this, keep the result. It is lazy: nothing is stored until an action runs.
clean = raw.filter(...).withColumn(...).cache() clean.count() # computes and caches clean.groupBy("a").count().show() # reads from cache clean.write.parquet("out") # reads from cache clean.unpersist()
CACHE TABLE clean AS SELECT ... FROM raw WHERE ...; SELECT a, COUNT(*) FROM clean GROUP BY a; UNCACHE TABLE clean;
Note the SQL difference: CACHE TABLE is eager by default and computes immediately; CACHE LAZY TABLE waits for the first use.
Storage levels
| Level | Where | Trade-off |
|---|---|---|
MEMORY_AND_DISK | Memory, spilling to disk | Default for DataFrame.cache(). Safe choice. |
MEMORY_ONLY | Memory only | Partitions that do not fit are recomputed when needed. |
DISK_ONLY | Local disk | Cheap on memory; reading back costs I/O. |
..._2 variants | Two copies on two executors | Survives losing an executor; doubles the space. |
OFF_HEAP | Off-heap memory | Needs off-heap memory configured. |
DataFrames are cached in Spark's compressed in-memory columnar format, so cached data is often smaller than the same rows as plain objects. Use df.persist(StorageLevel.DISK_ONLY) to choose a level.
When caching helps
- The DataFrame is used by two or more actions, or two branches of the same job.
- It is expensive to compute (joins, wide aggregations, slow sources such as JDBC or APIs) and much smaller than its inputs.
- Interactive exploration, where you query the same subset repeatedly.
When caching hurts
Caching something used once
Caching a plain read of Parquet
Caching huge DataFrames
Forgetting unpersist
Checking the cache
The Storage tab of the Spark UI lists cached DataFrames, the fraction cached and their size in memory and on disk. In the plan, cached data appears as InMemoryRelation / InMemoryTableScan.
checkpoint() is different: it writes the data to reliable storage and cuts the lineage, which helps with very long plans (iterative algorithms). localCheckpoint() does the same on executor storage, faster but not fault tolerant.Common mistakes
Expecting cache() to compute immediately
Caching after the last use
Caching inside a loop without unpersisting
Key takeaways
- cache() is lazy; the first action stores the data.
- DataFrame.cache() uses MEMORY_AND_DISK in a compressed columnar format.
- Cache only data reused across actions and costly to recompute.
- Unpersist when done; cached data competes with execution memory.
Check yourself
3 questions1. When is the data of df.cache() actually stored?
Show the answer
At the first action that computes df. cache only marks the DataFrame; the first action fills the cache.
2. Default storage level of DataFrame.cache()?
Show the answer
MEMORY_AND_DISK. DataFrames default to MEMORY_AND_DISK.
3. Why can caching a raw Parquet read slow later filtered queries?
Show the answer
Queries on the cache cannot skip files and row groups using Parquet statistics. Once cached, queries scan the cached data and lose file-level pushdown and pruning.
Go deeper
Transformations vs actionsSpark internals
The memory modelSpark internals
Jobs, stages and tasks
Primary sources: Caching data in memory · CACHE TABLE