In production, prefer an explicit schema when you read. Inference scans data (extra cost), can guess wrong types, and makes pipelines brittle when a dirty row appears.
Explicit schema example
from pyspark.sql.types import StructType, StructField, StringType, DoubleType, IntegerType, DateType
schema = StructType([
StructField("order_id", StringType(), False),
StructField("user_id", IntegerType(), True),
StructField("amount", DoubleType(), True),
StructField("order_date", DateType(), True),
StructField("status", StringType(), True),
])
orders = (
spark.read
.option("header", True)
.option("mode", "PERMISSIVE") # or FAILFAST / DROPMALFORMED
.schema(schema)
.csv("s3://lake/raw/orders/")
)Why not inferSchema alone
# Fine for exploration only
df = spark.read.option("header", True).option("inferSchema", True).csv(path)- Inference may promote IDs to integers incorrectly, or keep numbers as strings inconsistently.
- Schema-on-read contracts belong in code or a table metastore.
Diagram
CSV files --> reader (header + schema) --> DataFrame with known types
\-- bad rows --> _corrupt_record (PERMISSIVE)Interview add-on
Discuss mode (PERMISSIVE/FAILFAST), date formats, multiline JSON/CSV options, and enforcing schema at the bronze-to-silver boundary.