43. Data Quality: Null Count per Column
Difficulty: Medium · Topics: Null Handling, Filtering & Selection
Before loading customers, profile it: return a single row with one column per input column named <column>_nulls, holding how many rows have null in that column.
Write it so it works for any table, without typing column names.
Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.
Sample data
customers
| id | phone | city | |
|---|---|---|---|
| 1 | a@x.in | null | Pune |
| 2 | null | null | Goa |
| 3 | c@x.in | 99 | null |
Expected output
| id_nulls | email_nulls | phone_nulls | city_nulls |
|---|---|---|---|
| 0 | 1 | 2 | 1 |
Hints
Hint 1
customers.columns is a Python list of names.Hint 2
Build the expressions with a list comprehension and pass the list toselect.Hint 3
F.sum(F.col(c).isNull().cast("int")) counts nulls in column c.PySpark functions you'll practise
- sum
- isNull
- select
Related problems
- Sessionize a Clickstream · Hard · Window Functions
- Build an SCD Type 2 History · Hard · Window Functions
- Joining on Nullable Keys · Medium · Joins
- Price Valid at Order Time (Range Join) · Hard · Joins
- Reconcile Two Sources (Full Outer Join) · Hard · Joins
Browse
Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings
Difficulty: Easy · Medium · Hard · PySpark interview roadmap · All problems