Skip to content

26. Percent Rank and Cumulative Distribution

Difficulty: Medium · Topics: Window Functions

Within each dept, compute for every employee:

Round both to 2 decimals. Return name, dept, salary, pct_rank, cume_dist.

Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

employees

namedeptsalary
Anaeng100
Boeng120
Cyeng120
Dieng150
Edops80
Floops95

Expected output

namedeptsalarypct_rankcume_dist
Anaeng10000.25
Boeng1200.330.75
Cyeng1200.330.75
Dieng15011
Edops8000.5
Floops9511

Hints

Hint 1Both are window functions over Window.partitionBy("dept").orderBy("salary").
Hint 2percent_rank = (rank − 1) / (rows − 1); a one-person department gets 0.
Hint 3cume_dist counts ties (peers) together: equal salaries get the same value.

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · All problems