In PySpark DataFrames, `filter()` and `where()` are aliases. They do the same thing: keep rows that satisfy a boolean condition.
from pyspark.sql import functions as F
df.filter(F.col("amount") > 100)
df.where(F.col("amount") > 100)
df.filter("amount > 100") # SQL expression string also works
df.where("amount > 100")Why both exist
where reads more like SQL for analysts; filter matches RDD/functional naming. Teams usually pick one style and stick to it.
Diagram
DataFrame rows ---- condition ----> subset of rows
filter / where
(same)Related interview point
Do not confuse DataFrame filter with RDD filter, or with select (columns vs rows). Also remember NULL semantics: rows where the predicate is unknown are dropped.