Spark internals
20 in-depth lessons. How Spark executes your code, the configs that change it, and what is new in Spark 4.0. Every lesson ends with a quick check.
Everyone knows
- Driver, executors and the cluster manager · Everyone knows · 4 min read. Who plans the work, who does it, and who hands out the machines in a Spark application.
- Transformations vs actions · Everyone knows · 3 min read. Why nothing runs until you call an action, and what lazy evaluation buys Spark.
- Partitions · Everyone knows · 4 min read. The unit of parallelism in Spark: where partitions come from and why their number matters.
- Jobs, stages and tasks · Everyone knows · 3 min read. How one action becomes a DAG of stages split at shuffles, and how to read it in the Spark UI.
Good engineers know
- Narrow vs wide transformations · Good engineers know · 3 min read. Which operations move data across the network, and why the shuffle is the most expensive step.
- Broadcast joins · Good engineers know · 3 min read. Ship a small table to every executor and skip the shuffle, and when broadcasting backfires.
- Caching and persistence · Good engineers know · 3 min read. cache(), persist() and storage levels: when caching speeds a job up and when it hurts.
- Shuffle partitions · Good engineers know · 3 min read. The default of 200 shuffle partitions, and how to size them for your data.
- Adaptive Query Execution · Good engineers know · 4 min read. Spark re-plans a running query from real statistics: coalescing, join switching and skew splitting.
- The small file problem · Good engineers know · 6 min read. Why thousands of tiny files slow every read, how Spark jobs create them, and how to prevent and fix them.
Great engineers know
- Data skew and salting · Great engineers know · 5 min read. Why one task runs for an hour while the rest finish in seconds, and how salting spreads a hot key.
- AQE skew-join handling · Great engineers know · 3 min read. How Spark detects oversized partitions in sort-merge joins and splits them automatically.
- Catalyst and physical plans · Great engineers know · 3 min read. Read explain() output: parsed, analyzed and optimized plans, and the physical plan Spark runs.
- The memory model · Great engineers know · 4 min read. Execution vs storage memory, overhead, and what really causes spills and out-of-memory errors.
- Dynamic partition pruning · Great engineers know · 3 min read. Skip whole partitions of a fact table at runtime using a filter on the dimension it joins to.
- Spark Connect · Great engineers know · 3 min read. A thin client that talks to a remote Spark server over gRPC, and what it changes for applications.
- ANSI mode by default · Great engineers know · 3 min read. Spark 4.0 turns ANSI SQL on: overflows and invalid casts now raise errors instead of returning null.
- The VARIANT type · Great engineers know · 3 min read. Store and query semi-structured JSON efficiently without declaring a schema up front.
- SQL pipe syntax · Great engineers know · 3 min read. Write SQL top to bottom with the |> operator, in the order the query actually runs.
- Python Data Source API · Great engineers know · 3 min read. Write custom batch and streaming readers and writers for Spark in pure Python.