PySpark Data Lake
5 interview problems on partitioned folders, file formats and compression, small files, partition pruning and late data in the raw zone, from easy to hard.
Learn Data Lake
- Data lake storage: formats and layout · 4 min read. CSV, JSON, Avro, ORC and Parquet compared, compression, and how to lay out folders in a data lake.
- Lake, warehouse, lakehouse · 3 min read. What data lakes and warehouses each get right, and what the lakehouse borrows from both.
- Partitioning done right · 4 min read. When to partition a table, how to choose the column, and how over-partitioning backfires.
- Parquet and columnar storage · 4 min read. Row groups, column chunks and statistics: why columnar files let engines skip most of the data.
- The small file problem · 6 min read. Why thousands of tiny files slow every read, how Spark jobs create them, and how to prevent and fix them.
- Find Partitions That Need Compaction · Easy · Data Lake
- How Much Does Each Query Read? · Medium · Data Lake
- Inventory a Partitioned Folder · Medium · Data Lake
- Find Files Spark Cannot Split · Medium · Data Lake
- Late Data in the Raw Zone · Medium · Data Lake
Browse
Topics: Window Functions · Joins · Aggregations · Pivot, Unpivot & Rollup · Arrays · Null Handling · Conditional Logic · Dates · Filtering & Selection · Strings · Data Lake · Lakehouse · Spark Performance
Difficulty: Easy · Medium · Hard · PySpark interview roadmap · Learn · All problems