DStreams (Discretized Streams) are the older Spark Streaming API: a sequence of RDDs over time. You write RDD-style transforms. Limited optimizations, weaker event-time story.
Structured Streaming is the modern API: you write DataFrame/Dataset SQL-style code against an unbounded table. Catalyst optimizes it. First-class event time, watermarks, and richer sinks.
DStreams: Structured Streaming:
t0: RDD0 unbounded table "events"
t1: RDD1 .groupBy(window(col("ts"), "5 minutes"))
t2: RDD2 .count()
map/filter on RDDs SQL/DataFrame plan + triggersWhy Structured Streaming won
- One mental model with batch DataFrames
- Better exactly-once sink integrations
- Watermarking and late data handling
- DStreams are maintenance-mode / legacy in teaching materials
Interview tip: "Use Structured Streaming; DStreams are legacy RDD micro-batches."