Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Read CSV with explicit schema

PySpark · DataFrame API & I/O

Read CSV with explicit schema

Mediumpyspark-11
csvschemaiodataframe

Question

How do you read a CSV in PySpark with an explicit schema, and why is that better than inferSchema?

Solution

In production, prefer an explicit schema when you read. Inference scans data (extra cost), can guess wrong types, and makes pipelines brittle when a dirty row appears.

Explicit schema example

from pyspark.sql.types import StructType, StructField, StringType, DoubleType, IntegerType, DateType

schema = StructType([
    StructField("order_id", StringType(), False),
    StructField("user_id", IntegerType(), True),
    StructField("amount", DoubleType(), True),
    StructField("order_date", DateType(), True),
    StructField("status", StringType(), True),
])

orders = (
    spark.read
    .option("header", True)
    .option("mode", "PERMISSIVE")  # or FAILFAST / DROPMALFORMED
    .schema(schema)
    .csv("s3://lake/raw/orders/")
)

Why not inferSchema alone

# Fine for exploration only
df = spark.read.option("header", True).option("inferSchema", True).csv(path)
  • Inference may promote IDs to integers incorrectly, or keep numbers as strings inconsistently.
  • Schema-on-read contracts belong in code or a table metastore.

Diagram

CSV files --> reader (header + schema) --> DataFrame with known types
                 \-- bad rows --> _corrupt_record (PERMISSIVE)

Interview add-on

Discuss mode (PERMISSIVE/FAILFAST), date formats, multiline JSON/CSV options, and enforcing schema at the bronze-to-silver boundary.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext