Skip to content

42. Combine Batches with Schema Drift

Difficulty: Medium · Topics: Filtering & Selection

Two yearly extracts must be combined into one table. The 2024 extract has its columns in a different order and a new currency column that 2023 lacks.

Return all rows with columns id, name, amount, currency (in that order); 2023 rows have null currency.

Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

batch_2023

idnameamount
1Ana10
2Bo20

batch_2024

amountnameidcurrency
30Cy3INR
40Di4USD

Expected output

idnameamountcurrency
1Ana10null
2Bo20null
3Cy30INR
4Di40USD

Hints

Hint 1union matches columns by position, which would mix up name and amount here.
Hint 2unionByName matches by name; allowMissingColumns=True fills absent columns with null.
Hint 3Start from the 2023 table so its column order comes first.

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · All problems