33. Median and 90th Percentile
Difficulty: Medium · Topics: Aggregations
Per dept, compute median_salary (exact median: average of the two middle values for an even count) and p90_salary (percentile_approx at 0.9, which returns an actual salary from the data).
Missing salaries are ignored. Return dept, median_salary, p90_salary.
Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.
Sample data
employees
| dept | salary |
|---|---|
| eng | 100 |
| eng | 200 |
| eng | 300 |
| eng | 400 |
| ops | 50 |
| ops | 70 |
| ops | 90 |
Expected output
| dept | median_salary | p90_salary |
|---|---|---|
| eng | 250 | 400 |
| ops | 70 | 90 |
Hints
Hint 1
F.median("salary") (Spark 3.4+) interpolates between the middle values.Hint 2
F.percentile_approx("salary", 0.9) picks a real value: the smallest salary with at least 90% of values at or below it.Hint 3
Both skip nulls.PySpark functions you'll practise
- groupBy
- agg
- median
- percentile_approx
Related problems
- Average Salary by Department · Easy · Aggregations
- Word Count · Medium · Arrays
- Departments with Large Teams · Easy · Aggregations
- Three-Day Login Streak · Hard · Window Functions
- Pivot with Several Aggregations · Hard · Pivot, Unpivot & Rollup
Browse
Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings
Difficulty: Easy · Medium · Hard · PySpark interview roadmap · All problems