Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Parquet file structure

File formats & storage · Parquet Internals

Parquet file structure

Mediumfile-formats-21
parquetcolumnarstorage-internalsrow-groups

Question

What is inside a Parquet file?

Solution

A Parquet file is structured into horizontal row groups, which contain vertical column chunks, which in turn are divided into individual pages. At the end of the file sits a metadata file footer containing the table schema along with min/max statistics and null counts for every column chunk. Because the footer stores byte offsets for all structures, query engines read the footer first, evaluate pushdown filters against chunk statistics to prune entire row groups, and read only the byte ranges of needed columns.

Physical layout inside the file

+-------------------------------------------------------+
| Magic Number: 'PAR1'                                  |
+-------------------------------------------------------+
| Row Group 0 (e.g. 128 MB logical chunk of rows)       |
|   Column Chunk A (Dictionary Page, Data Pages 0..N)   |
|   Column Chunk B (Data Pages 0..N)                    |
+-------------------------------------------------------+
| Row Group 1 ...                                       |
+-------------------------------------------------------+
| File Footer (Metadata, Schema, Chunk Min/Max Stats)   |
| 4-byte Footer Length                                  |
| Magic Number: 'PAR1'                                  |
+-------------------------------------------------------+

Data in Parquet is organized hierarchically:

  • Row group: A horizontal partition of rows across all columns. A single row group typically buffers between 128 MB and 512 MB of uncompressed data in memory before being flushed to storage.
  • Column chunk: The data for a specific column within a row group, stored contiguously on disk so readers can scan that single column without touching adjacent fields.
  • Pages: The smallest unit of storage inside a column chunk, typically around 1 MB in size. Pages are compressed and encoded individually.
  • Dictionary page: An optional page preceding data pages that records unique values for dictionary encoding, allowing data pages to store compact integer dictionary indexes instead of repeating long strings.

Reader execution path

Engines like Spark, DuckDB, and Trino read Parquet files backwards:

  • The reader issues a small range request to read the final bytes of the file, parsing the 4-byte footer length and the Thrift metadata structures in the file footer.
  • The reader inspects the footer metadata to find the schema, projection column offsets, and min/max statistics for every column chunk in every row group.
  • If a query predicate specifies a filter, the reader compares the filter condition against the column chunk bounds. When the statistics prove no matching rows exist in a row group, the engine skips reading all byte ranges for that row group entirely.
  • Finally, the reader issues vectored read requests to fetch only the byte ranges corresponding to projected columns in surviving row groups.

Tuning row group sizing

Row group sizing involves a direct balance between memory overhead and pruning granularity:

  • Small row groups (such as 8 MB) inflate metadata volume, degrade columnar compression ratios, and increase footer parsing time on the query coordinator.
  • Excessively large row groups (such as 1 GB) require massive write buffers in executor memory and reduce data skipping opportunities, since wider row groups cover broader min/max value spans.

Targeting 128 MB to 512 MB per row group balances memory stability with effective predicate skipping.

PreviousNext