Skip to content

67. Find Hot Keys and Plan Salting

Difficulty: Hard · Topics: Joins, Aggregations, Filtering & Selection

Before a big join, you want to know which values of key in events are skewed. A key is hot when its row count is more than 3 times the median row count per key, where the median is percentile_approx(count, 0.5) over all keys.

For each hot key return key, rows and salt_buckets = the row count divided by the median, rounded up. Sort by rows descending, then key.

Row order matters for this problem. Your code is graded on 4 test cases, including hidden edge cases.

Sample data

events

event_idkey
1a
2a
3b
4b
5c
6c
7d
8hot
9hot
10hot
11hot
12hot
13hot
14hot

Expected output

keyrowssalt_buckets
hot74

Hints

Hint 1First count rows per key.
Hint 2Compute the median of those counts in a second aggregation, then bring it to every key with a cross join.
Hint 3F.ceil(F.col("rows") / F.col("median")) rounds up.

Learn the concepts

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems