PySpark joins mirror SQL joins. You specify the condition and a how type.
Common types
inner matching keys only left / left_outer right / right_outer full / full_outer left_semi left rows with a match on right (no right columns) left_anti left rows with NO match on right cross cartesian product (use carefully)
# Inner / left orders.join(customers, "customer_id", "inner") orders.join(customers, "customer_id", "left") # Existence checks without blowing up columns # Customers who placed at least one order customers.join(orders, "customer_id", "left_semi") # Customers with no orders customers.join(orders, "customer_id", "left_anti")
Diagram
Left Anti (A anti B): A rows whose key does not appear in B Left Semi (A semi B): A rows whose key appears in B (A columns only)
Why semi/anti beat some patterns
# Often worse: join then filter / distinct on right columns # Better: left_semi / left_anti for membership tests
Interview tip
Call out duplicate key blow-ups on inner/left joins, null-safe join keys (eqNullSafe), and broadcast hints for small dimensions.