Skip to content

33. Median and 90th Percentile

Difficulty: Medium · Topics: Aggregations

Per dept, compute median_salary (exact median: average of the two middle values for an even count) and p90_salary (percentile_approx at 0.9, which returns an actual salary from the data).

Missing salaries are ignored. Return dept, median_salary, p90_salary.

Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

employees

deptsalary
eng100
eng200
eng300
eng400
ops50
ops70
ops90

Expected output

deptmedian_salaryp90_salary
eng250400
ops7090

Hints

Hint 1F.median("salary") (Spark 3.4+) interpolates between the middle values.
Hint 2F.percentile_approx("salary", 0.9) picks a real value: the smallest salary with at least 90% of values at or below it.
Hint 3Both skip nulls.

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · All problems