Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Reading and writing JSON, CSV, Parquet

Python · Python for Data Pipelines

Reading and writing JSON, CSV, Parquet

Easypython-51
jsoncsvparquetfile-formatspandas

Question

How do you read and write JSON, CSV and Parquet in Python, and which would you choose for pipeline output?

Solution

Python can read all three easily. For pipeline output, choose Parquet in most cases, and use JSON Lines only where you need to append records one at a time or keep a flexible structure.

Reading and writing each

import csv, json
import pandas as pd

# CSV
with open("orders.csv", newline="") as f:
    rows = list(csv.DictReader(f))
df = pd.read_csv("orders.csv")

# JSON (one document) and JSON Lines (one object per line)
data = json.load(open("orders.json"))
df = pd.read_json("orders.jsonl", lines=True)

# Parquet
df.to_parquet("orders.parquet", compression="snappy")   # uses pyarrow
df = pd.read_parquet("orders.parquet", columns=["order_id", "amount"])

Comparing them

  • CSV is human readable and universal. It has no types, so everything is a string until you parse it, and dates and numbers get guessed. Quoting and delimiter problems (commas inside text, embedded newlines) are a constant source of bugs, and it compresses poorly and cannot read just some columns.
  • JSON handles nested structures and is the format of APIs. A single big JSON array has to be read as a whole. JSON Lines (one JSON object per line) solves that: you can stream it, append to it, and split it for parallel processing, so it suits logs and event streams.
  • Parquet is columnar and compressed, and stores the schema and types in the file. Readers can load only the columns they need, and skip data with filters using built-in statistics. A file is often several times smaller than the same CSV, and reads much faster.

What I would choose

For data that other pipeline steps or analysts will read (curated tables, intermediate results): Parquet. For raw event capture where records arrive one by one: JSON Lines, converted to Parquet later in batches. For handing a small file to a person with Excel: CSV. For an API response you save: JSON, as it came.

Practical points

Write Parquet in reasonably sized files (tens to hundreds of MB), not thousands of tiny ones, and partition large datasets into folders by date. Always write to a temporary name and rename when done, so readers do not see a half-written file.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext