Kafka Streams is a Java/Scala client library for building stream processing applications on Kafka. Your app is both a consumer and a producer, with higher-level DSL/Processor APIs.
What you get beyond a raw consumer
- Declarative operations:
filter,map,join,aggregate, windowing - State stores (local RocksDB) for aggregations and joins
- Fault-tolerant state via changelog topics
- Scaling by running more stream thread / app instances (partition assignment under the hood)
- Optional exactly-once processing guarantee
orders topic --> Kafka Streams app --> enriched / aggregates topic
|
local state store
(backed by changelog)Mental model
You write an application that embeds the Streams library. There is no separate Streams cluster. Scale-out = run more instances of your app with the same application.id.
vs plain consumer
| Plain consumer | Kafka Streams | |---|---| | You hand-code poll/process/produce | DSL for transforms | | You manage state yourself | Managed state stores | | You wire joins manually | Stream-table joins built-in |
vs Spark/Flink
Streams is library-embedded, lightweight for many Kafka-centric services. Spark/Flink are fuller processing platforms with different ops models.
Interview tip: "Kafka Streams = embeddable stream processing library on Kafka, with state and optional EOS; scale by adding app instances."