Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Schema inference vs explicit schema

PySpark · DataFrame API in Practice

Schema inference vs explicit schema

Easypyspark-74
schemainfer-schemastructtypedata-contracts

Question

Why should production jobs avoid inferSchema?

Solution

Production jobs should define the schema explicitly. inferSchema makes Spark read the data once just to guess types, and the guess can change from one run to the next.

What goes wrong with inference

  • Extra cost. For CSV and JSON, Spark scans the data (or a sample) before the real read. On a large input that is an extra pass over the files.
  • Unstable types. Suppose customer_id is numeric in yesterday's file, and today's file contains one value "C-17". Yesterday it was inferred as an integer, today it becomes a string. Your joins, casts and downstream tables now behave differently, with no error at the read step.
  • Wrong guesses. A column with leading zeros ("00123") is inferred as an integer and the zeros are gone. A column that is empty in the sampled rows becomes a string.
  • Silent drift. You never find out that the source changed, because Spark adjusts to it.

Define the schema

from pyspark.sql.types import StructType, StructField, LongType, StringType, DecimalType, TimestampType

schema = StructType([
    StructField("order_id", LongType(), False),
    StructField("customer_id", StringType(), True),
    StructField("amount", DecimalType(12, 2), True),
    StructField("created_at", TimestampType(), True),
])
df = spark.read.schema(schema).option("header", True).csv(path)

# a DDL string is shorter for simple cases
df = spark.read.schema("order_id BIGINT, customer_id STRING, amount DECIMAL(12,2), created_at TIMESTAMP").csv(path)

Why it is better

The read is faster. Types are the same every run. When a value does not fit, you get a NULL or a corrupt record that you can count and quarantine, instead of a changed schema. Schema drift becomes something you can detect: compare the incoming file header with the expected schema and alert.

When inference is fine

Exploring a new dataset in a notebook, or reading self-describing formats like Parquet and Avro, where the schema is stored in the file. For CSV or JSON, you can also reduce the cost with samplingRatio for JSON, but that does not fix unstable types. Keep schemas in code or a registry, version them, and review changes like any other code change.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext