Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. overwrite vs append

PySpark · DataFrame API & I/O

overwrite vs append

Mediumpyspark-20
write-modeoverwriteappendio

Question

What is the difference between overwrite and append write modes in PySpark?

Solution

Write mode controls how the sink handles existing data.

Modes

  • `append`: add new files next to existing ones. Does not remove old data.
  • `overwrite`: replace the target path (or matching partitions in dynamic partition overwrite).
  • Also: ignore (no-op if exists), error/errorifexists (default fail).
# Append new batch
df.write.mode("append").parquet("/lake/bronze/clicks")

# Replace entire dataset at path
df.write.mode("overwrite").parquet("/lake/silver/clicks_daily")

Diagram

append:     [old files] + [new files]
overwrite:  delete/replace target -> [new files only]

Production caveats

  • Blind overwrite of a shared path is dangerous in concurrent jobs.
  • Append without dedupe can double-count reprocessed batches; pair with idempotent keys or merge patterns (Delta/Iceberg MERGE).
  • For partitioned tables, dynamic partition overwrite rewrites only touched partitions instead of the whole table.
spark.conf.set("spark.sql.sources.partitionOverwriteMode", "dynamic")
(
    df.write
    .mode("overwrite")
    .partitionBy("dt")
    .parquet("/lake/silver/events")
)

Interview tip

Explain append vs overwrite, then mention lakehouse table formats for safer transactional writes.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext