Skip to content

19. Unpivot Wide Columns to Rows

Difficulty: Medium · Topics: Pivot, Unpivot & Rollup, Null Handling, Filtering & Selection

The budget table arrived in spreadsheet shape: one column per month (jan, feb, mar).

Turn it into a long table with columns dept, month (the column name, e.g. 'jan') and amount. Leave out months with no amount (null).

Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

budget

deptjanfebmar
Sales10012090
IT300null310

Expected output

deptmonthamount
Salesjan100
Salesfeb120
Salesmar90
ITjan300
ITmar310

Hints

Hint 1Spark 3.4+ has df.unpivot(ids, values, variableColumnName, valueColumnName) (alias melt).
Hint 2The ids stay as they are; each value column becomes one row per input row.
Hint 3Unpivot keeps nulls, so filter them afterwards.

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · All problems