Skip to content
Everyone knows 4 min read · All formats

Parquet and columnar storage

Row groups, column chunks and statistics: why columnar files let engines skip most of the data.

You will learn

  • How a Parquet file is laid out: row groups, column chunks, pages, footer
  • Why columnar storage reads less data
  • How statistics and dictionaries let engines skip data
  • How file and row-group size affect performance

Read first

Nothing. This lesson starts from scratch.

TL;DR Parquet stores data column by column within row groups, with min/max statistics for each column chunk in the file footer. An engine reads only the columns a query uses and skips row groups whose statistics rule out the filter.

Inside a Parquet file

File layout

part-00000.parquet
├── Row group 0  (e.g. 128 MB of rows)
│   ├── Column chunk: order_id   → pages
│   ├── Column chunk: country    → pages
│   └── Column chunk: amount     → pages
├── Row group 1
│   └── ...
└── Footer
    ├── Schema
    └── Per row group, per column: min, max, null count, offsets
    "PAR1"
  • Row group: a horizontal slice of rows. Writers commonly target around 128 MB.
  • Column chunk: the values of one column within a row group, stored together.
  • Page: the unit of encoding and compression inside a chunk, around 1 MB.
  • Footer: at the end of the file, the schema and statistics. Readers read it first.

Why columnar reads less

Analytic queries touch a few columns of wide tables. In a row format (CSV, JSON, Avro), reading one column means reading every byte of every row. In Parquet, the reader seeks straight to the column chunks it needs.

QueryRow format readsParquet reads
SELECT SUM(amount) on a 100-column tableAll 100 columns1 column (about 1% of the data, before compression)

Values of one column are also similar to each other, so they compress far better than mixed rows: dictionary encoding replaces repeated strings with small integers, run-length encoding collapses runs of the same value, and a general codec (Snappy by default in Spark, or ZSTD) compresses the result.

Skipping data with statistics

For a filter like amount > 10000, the reader checks each row group's max for amount in the footer. If the max is 9,500, the whole row group is skipped without reading it. This is predicate pushdown at the file level.

Query · the filter

SELECT * FROM orders WHERE order_date = '2025-03-14' on a file with four row groups.

Footer · the stats

Row group 0: dates Jan 1 to Feb 10. Row group 1: Feb 10 to Mar 9. Row group 2: Mar 9 to Apr 2. Row group 3: Apr 2 to May 1.

Skip · the result

Only row group 2 can contain March 14, so only it is read: 25% of the file. If the data had not been sorted by date, every row group would span all dates and nothing could be skipped.

That last point is the most important practical lesson: statistics help only when data is clustered by the filtered column. Sorting, Z-ordering and liquid clustering exist to make min/max ranges narrow.

Sizes that matter

  • Files of 128 MB to 1 GB keep per-file overhead low (see The small file problem).
  • Row groups around 128 MB balance skipping granularity against overhead; tiny row groups waste footer space and compress poorly.
  • Nested data (structs, arrays) is stored column by column too, using repetition and definition levels, so reading one nested field does not read the rest.
Note: Parquet is immutable: there is no way to update a row in place. Every update or delete means writing a new file. This is exactly why table formats exist.

Common mistakes

Expecting statistics to help on unsorted data

If every row group spans all values, nothing can be skipped.

Writing thousands of tiny Parquet files

Footer reads and opens dominate.

Selecting * from wide tables

Throws away the columnar advantage.

Key takeaways

  • Parquet stores columns separately inside row groups, with statistics in the footer.
  • Engines read only the needed columns and skip row groups using min/max values.
  • Skipping works only when data is clustered by the filtered column.
  • Parquet files are immutable; updates mean new files.

Check yourself

3 questions

1. Where are a Parquet file's column statistics stored?

Show the answer

In the footer. The footer holds the schema and per-row-group statistics.

2. A filter is on customer_id but the data is sorted by date. What does row-group skipping achieve for that filter?

Show the answer

Probably very little, since each row group spans many customer_ids. Min/max ranges are wide for unsorted columns, so few row groups can be ruled out.

3. Why does columnar data compress well?

Show the answer

Values of the same column are similar, suiting dictionary and run-length encoding. Homogeneous columns compress far better than mixed rows.

Go deeper

Primary sources: Apache Parquet documentation · Spark: Parquet files