AQE skew-join handling
How Spark detects oversized partitions in sort-merge joins and splits them automatically.
On this page
Show code in
Every code block on the page follows this.
You will learn
- How AQE decides that a join partition is skewed
- How it splits a skewed partition without breaking the join
- Which join types it supports
- How to tune and verify it
Read first
- Adaptive Query Execution · 4 min read
- Data skew and salting · 5 min read
Comfortable with these? Read on.
When a partition counts as skewed
Both conditions must hold:
| Config | Default | Condition |
|---|---|---|
spark.sql.adaptive.skewJoin.skewedPartitionFactor | 5 | Partition size is more than 5 × the median partition size |
spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes | 256MB | And larger than 256 MB |
spark.sql.adaptive.skewJoin.enabled | true | The feature is on (needs AQE on) |
The threshold stops AQE from splitting partitions that are relatively large but absolutely small: a 20 MB partition next to 1 MB ones is not worth the overhead.
How the split works
Before · one huge task
Split · by map output
Replicate · the other side
After · 190 tasks
Supported joins
| Join type | Which side can be split |
|---|---|
| Inner | Either or both |
| Left outer, left semi, left anti | Left side only |
| Right outer | Right side only |
| Full outer | Not supported |
The rule mirrors broadcasting: the side whose unmatched rows must be preserved can be split, because each piece is joined against the full other side. Splitting the non-preserved side would make unmatched rows ambiguous.
Tuning and verifying
- If a skewed partition is not split, check both conditions: on a cluster where every partition is large, 5 × median may never be reached; lower the factor.
- Some plans, such as a join whose output is immediately re-shuffled the same way, cannot be split without an extra shuffle.
spark.sql.adaptive.forceOptimizeSkewedJoin(3.3+) lets AQE add that shuffle. - In the final plan (SQL tab of the Spark UI), the join shows
skew=trueand anAQEShuffleReadnode reporting the number of skewed partitions and splits.
spark.conf.set("spark.sql.adaptive.skewJoin.skewedPartitionFactor", "3") spark.conf.set("spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes", "128MB")
SET spark.sql.adaptive.skewJoin.skewedPartitionFactor = 3; SET spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes = 128MB;
Common mistakes
Expecting it to fix skewed aggregations
Assuming full outer joins are covered
Leaving AQE off
Key takeaways
- A partition is skewed when it exceeds 5 × the median and 256 MB (defaults).
- AQE splits it and replicates the other side's matching partition.
- Only sort-merge-style joins; outer joins only on the preserved side; never full outer.
- Verify with skew=true in the final plan.
Check yourself
3 questions1. With defaults, median partition 100 MB: is a 400 MB partition split?
Show the answer
No: it is not more than 5× the median. 400 MB is 4× the median, below the factor of 5.
2. What does AQE do to the non-skewed side's matching partition?
Show the answer
Reads it once per split of the skewed side. Each piece needs all matching rows, so that partition is read by every split task.
3. In a left outer join, which side can AQE split?
Show the answer
Left. The preserved (left) side can be split; each piece joins against the full right partition.
Practice it
Interview problems that use this: write the PySpark, run it, and get graded on hidden tests.
Go deeper
Adaptive Query ExecutionSpark internals
Data skew and saltingSpark internals
Broadcast joins
Primary sources: AQE: optimizing skew join