Learn · streaming
Processing events as they arrive instead of waiting for tonight's batch.
Why streaming exists, how message queues work, partitions, offsets, consumer groups, late data, and replay. Simulated topics (no broker in this tab).
Core foundational modules are free. Advanced production modules need Pro.
Playable walkthroughs: watch the system move, predict the next step, stamp a memory seal, then practice. Completing a Trace counts toward readiness.
A fraud team needs to block a suspicious card in seconds. A nightly batch job tells them tomorrow, which is useless. So the data has to move continuously: every transaction, as it happens. That changes everything about how you handle it. There is no end of file, so you never know if you have seen everything. A network hiccup means a message might arrive twice. An event from a phone with bad signal might arrive twenty minutes late, after you already reported the total. Streaming systems are built around those three facts.
Fewer jobs are pure streaming than job ads suggest, but the vocabulary shows up constantly: topics, partitions, offsets, consumer groups, exactly once, watermarks. Interviews ask about it even for batch roles, because the reasoning about duplicates and late data applies everywhere.
Streaming is best understood as a contrast with batch. If you have never built a batch job, do that first.
Producers and consumers here are Python, simulated in the browser. No broker to install.
6 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Start with when streaming is genuinely worth it, because it costs more to build and operate than batch. Most data does not need it.
Batch versus StreamingFree
0/3
When batch is enough, when you need real-time, common architectures, and why exactly-once delivery is hard.
By the end of this stage
You can decide whether a requirement actually needs streaming, and say why.
Connect every producer directly to every consumer and you get a mess that breaks whenever one side is down. A log in the middle decouples them, and that single idea is what Kafka sells.
Why QueuesFree
0/2
Why a message queue exists, then how FakeKafkaTopic splits keys across partitions.
By the end of this stage
You can explain what a queue buys you, in terms of failure and speed mismatch between producer and consumer.
This is the mechanical core. Ordering, parallelism, and how a consumer knows where it left off after a crash all come from these three ideas.
Offsets and Groups
0/2
A bookmark per partition, then two consumers in one group splitting the aisles.
By the end of this stage
You can explain how work is split across consumers, what ordering guarantee you actually get, and what happens when a consumer dies mid-batch.
Where people get stuck
Ordering is per partition, not global. Most streaming misunderstandings trace back to missing that.
At least once, at most once, exactly once, and event time versus processing time. These are the topics that separate people who have run a stream from people who have read about one.
Delivery and Time
0/3
At-least-once vs a commit, late event_time, then a watermark that is not the PySpark exercise.
By the end of this stage
You can explain what a watermark is for, and what tradeoff you make when you choose how long to wait for late events.
Where people get stuck
Exactly once is widely misunderstood. Learn what it actually guarantees and where the boundary is.
Real systems are both. The interesting question is how to keep a streaming view and a batch view from disagreeing.
Stream-batch
0/2
Batch is a bounded stream. Replay is seek-and-reread on the same FakeKafkaTopic.
By the end of this stage
You can describe how a real architecture serves fresh data and correct historical data at the same time.
Put the pieces together on a realistic event flow with duplicates and late arrivals.
Capstone
0/1
Produce three FakeKafkaTopic records, drop the late one, keep ORD-2 and ORD-3.
By the end of this stage
You can design a consumer that is safe to restart and honest about lateness.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 6 sessions
1 hour a day
about 3 sessions
4 hours a weekend day
about 1 session
Topics here are simulated, so there is no broker to run. Spend your time on offsets and delivery guarantees, because those two modules carry almost all the interview weight and almost all the real-world pain.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Assuming exactly once means duplicates are impossible everywhere. | It is a guarantee about a specific boundary. Your sink still usually needs idempotent writes. |
| Committing the offset before the work is done. | A crash then loses the message silently. Commit after the write, and make the write safe to repeat. |
| Choosing streaming because it sounds more advanced. | Streaming doubles operational burden. If the business can wait an hour, batch is the better engineering decision. |