Skip to content

47. Inventory a Partitioned Folder

Difficulty: Medium · Topics: Data Lake, Aggregations, Strings

files is a listing of a Parquet table in a data lake: path and size_mb. The table is partitioned Hive-style, for example s3://lake/orders/order_date=2025-03-01/country=IN/part-00000.parquet.

For every partition return order_date, country, data_files and total_mb (rounded to 1 decimal). Count only data files: paths ending in .parquet where no folder or file name starts with _ or . (that excludes _SUCCESS markers, .crc checksums and unfinished _temporary writes). Spark writes a null partition value as __HIVE_DEFAULT_PARTITION__: return it as a real null.

Row order does not matter; column names must match. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

files

pathsize_mb
s3://lake/orders/order_date=2025-03-01/country=IN/part-00000.parquet120.5
s3://lake/orders/order_date=2025-03-01/country=IN/part-00001.parquet98.25
s3://lake/orders/order_date=2025-03-01/country=IN/.part-00000.parquet.crc0.01
s3://lake/orders/order_date=2025-03-01/country=US/part-00002.parquet64
s3://lake/orders/order_date=2025-03-01/country=__HIVE_DEFAULT_PARTITION__/part-00003.parquet2.4
s3://lake/orders/_SUCCESS0
s3://lake/orders/order_date=2025-03-02/country=IN/part-00000.parquet110

Expected output

order_datecountrydata_filestotal_mb
2025-03-01IN2218.8
2025-03-01US164
2025-03-01null12.4
2025-03-02IN1110

Hints

Hint 1Pull a partition value out of the path with F.regexp_extract("path", r"/order_date=([^/]+)/", 1).
Hint 2A name starting with _ or . always follows a slash: F.col("path").rlike("/[_.]") finds them.
Hint 3Turn the placeholder into null with F.when(col == "__HIVE_DEFAULT_PARTITION__", None).otherwise(col), then group and aggregate.

Learn the concepts

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings · Data Lake · Lakehouse · Spark Performance

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems