Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Splittable vs non-splittable files

File formats & storage · Parquet Internals

Splittable vs non-splittable files

Mediumfile-formats-24
splittableparquetavroexecutionparallelism

Question

What does it mean for a file to be splittable, and why does it matter?

Solution

A file is splittable if a distributed processing engine like Spark or Trino can independently read arbitrary byte ranges (splits) of that single file across parallel tasks without reading from the beginning. This matters because if a large file cannot be split, a single executor core must process the entire file sequentially, creating a severe straggler task that destroys pipeline parallelism. Columnar and binary formats like Parquet, ORC, and Avro are splittable by design, whereas stream-compressed text formats like GZIP-compressed CSV are non-splittable.

Splittable file mechanics

In a distributed framework, an input split specifies a byte offset and length, such as bytes 0 to 128 MB for task 1, and bytes 128 MB to 256 MB for task 2. For an engine to process a split, the file format must provide synchronization markers or internal metadata boundaries:

  • Parquet and ORC store self-contained row groups or stripes; the engine reads the metadata footer once, discovers exact row group byte offsets, and schedules separate worker tasks for each row group.
  • Avro inserts 16-byte magic sync markers at regular intervals between data blocks, allowing a worker that seeks to an arbitrary offset to scan forward until it finds a sync marker and begin decoding records.
  • Uncompressed CSV is splittable because workers can seek to a byte offset and discard characters until the first newline byte, resuming clean record parsing from that line onward.

The compressed text bottleneck

The catastrophe occurs when someone compresses large text files with GZIP. Because GZIP maintains a sliding 32 KB window and dynamic Huffman codes across the entire byte stream, worker tasks cannot jump into byte offset 500 MB without knowing the compression state from byte 0. If an upstream team dumps a 20 GB .csv.gz file into S3, Spark assigns exactly one task with one core to read it. Even on a 500-core cluster, 499 cores sit idle while a single executor spends an hour decompressing the file.

Practical remedies for non-splittable data

When you encounter non-splittable compressed files, apply two immediate remedies:

  • Convert the data immediately into Parquet or Snappy-compressed Avro at the ingestion boundary so downstream jobs read splittable blocks.
  • If upstream systems must deliver GZIP CSV, instruct the sender to partition the export into many small files (such as 100 MB each) so that cluster tasks achieve parallelism across individual file handles.
PreviousNext