JSON Lines (newline-delimited JSON or NDJSON) stores one complete, self-contained JSON object per line separated by a newline character, whereas a JSON array wraps all records inside a single top-level bracket. This makes NDJSON streamable and splittable across distributed worker tasks because any worker can seek to an arbitrary byte offset and begin reading at the next newline. In contrast, a standard JSON array requires the parser to load and scan the entire file structure as a single document tree, which bottlenecks memory and eliminates read parallelism.
Line streaming vs document tree parsing
JSON Array (Monolithic):
[
{"id": 1, "status": "active"},
{"id": 2, "status": "pending"}
] --> Parser must hold document context across all lines
JSON Lines / NDJSON (Streaming):
{"id": 1, "status": "active"} --> Line 1 (Independent record)
{"id": 2, "status": "pending"} --> Line 2 (Independent record)The structural advantages of NDJSON become evident during pipeline execution:
- Streamable ingestion: Log producers and Kafka connectors can continuously append individual lines to an S3 object or file stream without having to open, parse, rewrite, and close a terminal array bracket.
- Distributed splitting: Spark and Trino can divide a 5 GB uncompressed NDJSON file into forty 128 MB input splits, allowing forty cores to parse lines concurrently.
- Error isolation: If a single line contains corrupt JSON syntax, a streaming reader can quarantine that specific bad line using a corrupted record column rather than aborting the entire multi-gigabyte dataset.
Spark memory impact with multiLine
When processing a standard JSON array in Spark, you must set the multiLine option to true:
- This option forces Spark to disable file splitting entirely, reading the file on a single executor core into memory.
- If the array file is several gigabytes, this frequently causes OutOfMemoryError crashes on the driver or executor.
- Memory pressure escalates because the JSON tree parser must instantiate the entire object hierarchy before emitting rows.
Ingestion to columnar storage
Even though NDJSON solves file splitting and streaming appends, it remains a row-oriented, text-heavy format that stores repetitive column keys for every row. In production systems, ingest NDJSON at the edge, validate its schema, and immediately compact it into columnar Parquet or table format tables for analytical consumption.