Parquet and columnar storage
Row groups, column chunks and statistics: why columnar files let engines skip most of the data.
On this page
You will learn
- How a Parquet file is laid out: row groups, column chunks, pages, footer
- Why columnar storage reads less data
- How statistics and dictionaries let engines skip data
- How file and row-group size affect performance
Read first
Nothing. This lesson starts from scratch.
Inside a Parquet file
File layout
part-00000.parquet
├── Row group 0 (e.g. 128 MB of rows)
│ ├── Column chunk: order_id → pages
│ ├── Column chunk: country → pages
│ └── Column chunk: amount → pages
├── Row group 1
│ └── ...
└── Footer
├── Schema
└── Per row group, per column: min, max, null count, offsets
"PAR1"- Row group: a horizontal slice of rows. Writers commonly target around 128 MB.
- Column chunk: the values of one column within a row group, stored together.
- Page: the unit of encoding and compression inside a chunk, around 1 MB.
- Footer: at the end of the file, the schema and statistics. Readers read it first.
Why columnar reads less
Analytic queries touch a few columns of wide tables. In a row format (CSV, JSON, Avro), reading one column means reading every byte of every row. In Parquet, the reader seeks straight to the column chunks it needs.
| Query | Row format reads | Parquet reads |
|---|---|---|
SELECT SUM(amount) on a 100-column table | All 100 columns | 1 column (about 1% of the data, before compression) |
Values of one column are also similar to each other, so they compress far better than mixed rows: dictionary encoding replaces repeated strings with small integers, run-length encoding collapses runs of the same value, and a general codec (Snappy by default in Spark, or ZSTD) compresses the result.
Skipping data with statistics
For a filter like amount > 10000, the reader checks each row group's max for amount in the footer. If the max is 9,500, the whole row group is skipped without reading it. This is predicate pushdown at the file level.
Query · the filter
SELECT * FROM orders WHERE order_date = '2025-03-14' on a file with four row groups.Footer · the stats
Skip · the result
That last point is the most important practical lesson: statistics help only when data is clustered by the filtered column. Sorting, Z-ordering and liquid clustering exist to make min/max ranges narrow.
Sizes that matter
- Files of 128 MB to 1 GB keep per-file overhead low (see The small file problem).
- Row groups around 128 MB balance skipping granularity against overhead; tiny row groups waste footer space and compress poorly.
- Nested data (structs, arrays) is stored column by column too, using repetition and definition levels, so reading one nested field does not read the rest.
Common mistakes
Expecting statistics to help on unsorted data
Writing thousands of tiny Parquet files
Selecting * from wide tables
Key takeaways
- Parquet stores columns separately inside row groups, with statistics in the footer.
- Engines read only the needed columns and skip row groups using min/max values.
- Skipping works only when data is clustered by the filtered column.
- Parquet files are immutable; updates mean new files.
Check yourself
3 questions1. Where are a Parquet file's column statistics stored?
Show the answer
In the footer. The footer holds the schema and per-row-group statistics.
2. A filter is on customer_id but the data is sorted by date. What does row-group skipping achieve for that filter?
Show the answer
Probably very little, since each row group spans many customer_ids. Min/max ranges are wide for unsorted columns, so few row groups can be ruled out.
3. Why does columnar data compress well?
Show the answer
Values of the same column are similar, suiting dictionary and run-length encoding. Homogeneous columns compress far better than mixed rows.
Go deeper
Open table formatsLakehouse
Compaction, OPTIMIZE and Z-orderSpark internals
The small file problem
Primary sources: Apache Parquet documentation · Spark: Parquet files