Skip to content

48. Find Files Spark Cannot Split

Difficulty: Medium · Topics: Data Lake, Aggregations, Conditional Logic

files lists the files in each zone of a data lake: zone, path and size_mb.

A gzip-compressed file (path ending in .gz) cannot be split, so Spark reads it with a single task. Such a file is slow when it is larger than 128 MB. For each zone return files (all files), gz_files, slow_files and slow_mb (total size of slow files, 0 when there are none). Sort by slow_mb descending, then zone.

Row order matters for this problem. Your code is graded on 3 test cases, including hidden edge cases.

Sample data

files

zonepathsize_mb
raws3://lake/raw/clicks/2025-03-01.json.gz2100
raws3://lake/raw/clicks/2025-03-02.json.gz90
raws3://lake/raw/orders/orders.csv.gz400
raws3://lake/raw/orders/orders.csv700
silvers3://lake/silver/clicks/part-0.snappy.parquet256
silvers3://lake/silver/clicks/part-1.snappy.parquet250
golds3://lake/gold/daily.csv.gz129

Expected output

zonefilesgz_filesslow_filesslow_mb
raw4322500
gold111129
silver2000

Hints

Hint 1Define the conditions once: gz = F.col("path").endswith(".gz") and slow = gz & (F.col("size_mb") > 128).
Hint 2Count matches inside one groupBy with F.sum(F.when(cond, 1).otherwise(0)).
Hint 3Use otherwise(0) for the size too, so a zone with no slow files sums to 0 instead of null.

Learn the concepts

PySpark functions you'll practise

Related problems

Browse

Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings · Data Lake · Lakehouse · Spark Performance

Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems