A stream-static join combines each micro-batch of a stream with a fixed table, and keeps no state. A stream-stream join has to remember rows from both streams so they can find each other later, which needs watermarks to keep state from growing forever.
Stream-static
enriched = (clicks_stream
.join(dim_users, "user_id", "left"))Each micro-batch of clicks_stream joins with dim_users. If the static side is a Delta table, Spark can re-read it for every batch, so updates to the dimension show up, while a plain file source may be read once. There is no state. The join is as cheap as a batch join, and you can use a broadcast join to avoid shuffling.
The limit is that the stream side drives it. A change in the static table does not produce new output for old events.
Stream-stream
Both inputs are unbounded, such as ad impressions and ad clicks. A click can arrive long after its impression. Spark stores the rows from each side in its state store, and every new row on one side is matched against the stored rows from the other side.
imp = impressions.withWatermark("imp_ts", "2 hours")
clk = clicks.withWatermark("click_ts", "3 hours")
joined = imp.join(
clk,
F.expr("""imp_id = click_imp_id AND
click_ts BETWEEN imp_ts AND imp_ts + INTERVAL 1 HOUR"""),
"inner")The two parts that matter are the watermarks and the time condition. Together they tell Spark when a stored row can no longer match anything, so it can drop it. Without them, Spark keeps every row from both sides forever, and state grows until the job fails.
Outer joins
For a left outer join on two streams, Spark can only emit an unmatched row (with NULLs) after it is sure that no match can arrive, which is when the watermark passes the time limit. So outer joins require watermarks and time constraints, and results for unmatched rows are delayed.
What to say
State the cost difference: stream-static is cheap and stateless, stream-stream is stateful and needs watermark and time bounds. Mention monitoring state size in the streaming progress metrics.