PySpark Spark Performance
2 interview problems on data skew and salting, join strategies and broadcast joins, from easy to hard.
Learn Spark Performance
- Data skew and salting · 5 min read. Why one task runs for an hour while the rest finish in seconds, and how salting spreads a hot key.
- Broadcast joins · 3 min read. Ship a small table to every executor and skip the shuffle, and when broadcasting backfires.
- Adaptive Query Execution · 4 min read. Spark re-plans a running query from real statistics: coalescing, join switching and skew splitting.
- AQE skew-join handling · 3 min read. How Spark detects oversized partitions in sort-merge joins and splits them automatically.
- Bucketing · 3 min read. Pre-shuffle tables into buckets so joins and aggregations on the key skip the shuffle.
- Pick the Join Strategy · Medium · Spark Performance
- Find Hot Keys and Plan Salting · Hard · Spark Performance
Browse
Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings · Data Lake · Lakehouse · Spark Performance
Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems