Production jobs should define the schema explicitly. inferSchema makes Spark read the data once just to guess types, and the guess can change from one run to the next.
What goes wrong with inference
- Extra cost. For CSV and JSON, Spark scans the data (or a sample) before the real read. On a large input that is an extra pass over the files.
- Unstable types. Suppose
customer_idis numeric in yesterday's file, and today's file contains one value"C-17". Yesterday it was inferred as an integer, today it becomes a string. Your joins, casts and downstream tables now behave differently, with no error at the read step. - Wrong guesses. A column with leading zeros (
"00123") is inferred as an integer and the zeros are gone. A column that is empty in the sampled rows becomes a string. - Silent drift. You never find out that the source changed, because Spark adjusts to it.
Define the schema
from pyspark.sql.types import StructType, StructField, LongType, StringType, DecimalType, TimestampType
schema = StructType([
StructField("order_id", LongType(), False),
StructField("customer_id", StringType(), True),
StructField("amount", DecimalType(12, 2), True),
StructField("created_at", TimestampType(), True),
])
df = spark.read.schema(schema).option("header", True).csv(path)
# a DDL string is shorter for simple cases
df = spark.read.schema("order_id BIGINT, customer_id STRING, amount DECIMAL(12,2), created_at TIMESTAMP").csv(path)Why it is better
The read is faster. Types are the same every run. When a value does not fit, you get a NULL or a corrupt record that you can count and quarantine, instead of a changed schema. Schema drift becomes something you can detect: compare the incoming file header with the expected schema and alert.
When inference is fine
Exploring a new dataset in a notebook, or reading self-describing formats like Parquet and Avro, where the schema is stored in the file. For CSV or JSON, you can also reduce the cost with samplingRatio for JSON, but that does not fix unstable types. Keep schemas in code or a registry, version them, and review changes like any other code change.