Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. JSON lines vs JSON array

File formats & storage · Parquet Internals

JSON lines vs JSON array

Easyfile-formats-26
jsonndjsonstreamingsparkparsing

Question

Why are JSON lines (NDJSON) better than a single JSON array for big data?

Solution

JSON Lines (newline-delimited JSON or NDJSON) stores one complete, self-contained JSON object per line separated by a newline character, whereas a JSON array wraps all records inside a single top-level bracket. This makes NDJSON streamable and splittable across distributed worker tasks because any worker can seek to an arbitrary byte offset and begin reading at the next newline. In contrast, a standard JSON array requires the parser to load and scan the entire file structure as a single document tree, which bottlenecks memory and eliminates read parallelism.

Line streaming vs document tree parsing

JSON Array (Monolithic):
[
  {"id": 1, "status": "active"},
  {"id": 2, "status": "pending"}
]  --> Parser must hold document context across all lines

JSON Lines / NDJSON (Streaming):
{"id": 1, "status": "active"}    --> Line 1 (Independent record)
{"id": 2, "status": "pending"}   --> Line 2 (Independent record)

The structural advantages of NDJSON become evident during pipeline execution:

  • Streamable ingestion: Log producers and Kafka connectors can continuously append individual lines to an S3 object or file stream without having to open, parse, rewrite, and close a terminal array bracket.
  • Distributed splitting: Spark and Trino can divide a 5 GB uncompressed NDJSON file into forty 128 MB input splits, allowing forty cores to parse lines concurrently.
  • Error isolation: If a single line contains corrupt JSON syntax, a streaming reader can quarantine that specific bad line using a corrupted record column rather than aborting the entire multi-gigabyte dataset.

Spark memory impact with multiLine

When processing a standard JSON array in Spark, you must set the multiLine option to true:

  • This option forces Spark to disable file splitting entirely, reading the file on a single executor core into memory.
  • If the array file is several gigabytes, this frequently causes OutOfMemoryError crashes on the driver or executor.
  • Memory pressure escalates because the JSON tree parser must instantiate the entire object hierarchy before emitting rows.

Ingestion to columnar storage

Even though NDJSON solves file splitting and streaming appends, it remains a row-oriented, text-heavy format that stores repetitive column keys for every row. In production systems, ingest NDJSON at the edge, validate its schema, and immediately compact it into columnar Parquet or table format tables for analytical consumption.

PreviousNext