With the default settings, overwriting a partitioned table replaces the whole table, not just the partitions you wrote. This is a classic way to lose data, so know which mode you are in.
Static overwrite (the default)
df_today.write.mode("overwrite").partitionBy("order_date").parquet("/lake/orders")If df_today only has data for 2025-03-01, Spark first deletes every partition under /lake/orders, then writes the new one. All previous days are gone. It surprises people who expected "replace just this day".
Dynamic partition overwrite
spark.conf.set("spark.sql.sources.partitionOverwriteMode", "dynamic")
df_today.write.mode("overwrite").partitionBy("order_date").parquet("/lake/orders")Now Spark only replaces partitions that appear in the DataFrame. Other days are untouched. If you rerun the job for the same day, that day is replaced by the new output, not duplicated. That is what makes daily reruns idempotent.
A caution
If the DataFrame accidentally contains rows for more dates than you intended (a bad filter, a late record from last month), those partitions are replaced as well, with only the rows in the DataFrame. Filter the input to the run date and assert it.
Table formats
Delta Lake uses a more explicit way: replaceWhere.
(df_today.write.format("delta").mode("overwrite")
.option("replaceWhere", "order_date = '2025-03-01'")
.save("/lake/orders"))It overwrites only rows matching the predicate, and fails if the new data has rows outside it, which protects you from the accident above. Delta and Iceberg also support MERGE, and Iceberg has its own dynamic overwrite behaviour. Look up your table format rather than assuming the Parquet behaviour.
What to say in an interview
Mention that the default is destructive, name the dynamic setting, and connect it to idempotent reruns. If you use a table format, say you prefer replaceWhere or MERGE because they are transactional: a failed job does not leave half the partitions deleted.