Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. dropDuplicates vs distinct

PySpark · DataFrame API & I/O

dropDuplicates vs distinct

Mediumpyspark-16
dedupedistinctdropduplicates

Question

What is the difference between dropDuplicates and distinct in PySpark?

Solution

Both remove duplicate rows, but `dropDuplicates` (alias `drop_duplicates`) can target a subset of columns, while `distinct()` dedupes on all columns.

# All columns
df.distinct()
df.dropDuplicates()          # same idea when no subset given

# Subset of columns (keep one row per key)
df.dropDuplicates(["user_id", "order_date"])

Behavior notes

distinct()
  -> equivalent to SELECT DISTINCT * ; shuffles

dropDuplicates([cols])
  -> one row per unique combination of cols
  -> which row is kept among ties is not a stable "business rule"
     unless you define order (window + filter)

Deterministic dedupe pattern

from pyspark.sql import functions as F
from pyspark.sql.window import Window

w = Window.partitionBy("user_id", "event_id").orderBy(F.col("ingest_ts").desc())
deduped = (
    df.withColumn("rn", F.row_number().over(w))
      .filter(F.col("rn") == 1)
      .drop("rn")
)

Interview tip

Say when you need "latest record wins" you do not rely on dropDuplicates alone; you use a window ordered by a timestamp.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext