48. Find Files Spark Cannot Split
Difficulty: Medium · Topics: Data Lake, Aggregations, Conditional Logic
files lists the files in each zone of a data lake: zone, path and size_mb.
A gzip-compressed file (path ending in .gz) cannot be split, so Spark reads it with a single task. Such a file is slow when it is larger than 128 MB. For each zone return files (all files), gz_files, slow_files and slow_mb (total size of slow files, 0 when there are none). Sort by slow_mb descending, then zone.
Row order matters for this problem. Your code is graded on 3 test cases, including hidden edge cases.
Sample data
files
| zone | path | size_mb |
|---|---|---|
| raw | s3://lake/raw/clicks/2025-03-01.json.gz | 2100 |
| raw | s3://lake/raw/clicks/2025-03-02.json.gz | 90 |
| raw | s3://lake/raw/orders/orders.csv.gz | 400 |
| raw | s3://lake/raw/orders/orders.csv | 700 |
| silver | s3://lake/silver/clicks/part-0.snappy.parquet | 256 |
| silver | s3://lake/silver/clicks/part-1.snappy.parquet | 250 |
| gold | s3://lake/gold/daily.csv.gz | 129 |
Expected output
| zone | files | gz_files | slow_files | slow_mb |
|---|---|---|---|---|
| raw | 4 | 3 | 2 | 2500 |
| gold | 1 | 1 | 1 | 129 |
| silver | 2 | 0 | 0 | 0 |
Hints
Hint 1
Define the conditions once:gz = F.col("path").endswith(".gz") and slow = gz & (F.col("size_mb") > 128).Hint 2
Count matches inside one groupBy withF.sum(F.when(cond, 1).otherwise(0)).Hint 3
Useotherwise(0) for the size too, so a zone with no slow files sums to 0 instead of null.Learn the concepts
- The small file problem · 6 min read. Why thousands of tiny files slow every read, how Spark jobs create them, and how to prevent and fix them.
- Data lake storage: formats and layout · 4 min read. CSV, JSON, Avro, ORC and Parquet compared, compression, and how to lay out folders in a data lake.
- when / otherwise · 3 min read. CASE WHEN logic inside a column: first match wins, and a missing otherwise means null.
PySpark functions you'll practise
- groupBy
- agg
- sum
- count
- when / otherwise
- orderBy
Related problems
- Late Data in the Raw Zone · Medium · Data Lake
- How Much Does Each Query Read? · Medium · Data Lake
- Inventory a Partitioned Folder · Medium · Data Lake
- Find Partitions That Need Compaction · Easy · Data Lake
- Conditional Aggregation · Medium · Aggregations
Browse
Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings · Data Lake · Lakehouse · Spark Performance
Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems