Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Learn
  3. Streaming & Message Queues

Learn · streaming

Streaming & Message Queues

Processing events as they arrive instead of waiting for tonight's batch.

Why streaming exists, how message queues work, partitions, offsets, consumer groups, late data, and replay. Simulated topics (no broker in this tab).

13 lessons6 modules6 stages2h 54m
Start lesson 1Batch versus streaming

Core foundational modules are free. Advanced production modules need Pro.

Concept Traces in this track

Playable walkthroughs: watch the system move, predict the next step, stamp a memory seal, then practice. Completing a Trace counts toward readiness.

  • Pro Trace

    Offsets and consumer groups

    The log keeps the tape. Your offset is the bookmark.

    Opens in Offsets and commits

  • Pro Trace

    Batch vs stream

    The SLA picks the clock. Tools come after.

    Opens in Batch versus streaming

Why this track exists

A fraud team needs to block a suspicious card in seconds. A nightly batch job tells them tomorrow, which is useless. So the data has to move continuously: every transaction, as it happens. That changes everything about how you handle it. There is no end of file, so you never know if you have seen everything. A network hiccup means a message might arrive twice. An event from a phone with bad signal might arrive twenty minutes late, after you already reported the total. Streaming systems are built around those three facts.

Fewer jobs are pure streaming than job ads suggest, but the vocabulary shows up constantly: topics, partitions, offsets, consumer groups, exactly once, watermarks. Interviews ask about it even for batch roles, because the reasoning about duplicates and late data applies everywhere.

What you need before starting

  • Batch pipeline understanding

    Streaming is best understood as a contrast with batch. If you have never built a batch job, do that first.

  • Python

    Producers and consumers here are Python, simulated in the browser. No broker to install.

The roadmap

6 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.

01Batch versus streaming

Start with when streaming is genuinely worth it, because it costs more to build and operate than batch. Most data does not need it.

Batch versus StreamingFree

0/3

When batch is enough, when you need real-time, common architectures, and why exactly-once delivery is hard.

  1. Batch versus streaming12m
  2. Streaming architectures12m
  3. Exactly-once semantics14m

By the end of this stage

You can decide whether a requirement actually needs streaming, and say why.

02Why a queue in the middle

Connect every producer directly to every consumer and you get a mess that breaks whenever one side is down. A log in the middle decouples them, and that single idea is what Kafka sells.

Why QueuesFree

0/2

Why a message queue exists, then how FakeKafkaTopic splits keys across partitions.

  1. Why a queue12m
  2. Topics, partitions, keys14m

By the end of this stage

You can explain what a queue buys you, in terms of failure and speed mismatch between producer and consumer.

03Partitions, offsets, and consumer groups

This is the mechanical core. Ordering, parallelism, and how a consumer knows where it left off after a crash all come from these three ideas.

Offsets and Groups

0/2

A bookmark per partition, then two consumers in one group splitting the aisles.

  1. Offsets and commits14m
  2. Consumer groups14m

By the end of this stage

You can explain how work is split across consumers, what ordering guarantee you actually get, and what happens when a consumer dies mid-batch.

Where people get stuck

Ordering is per partition, not global. Most streaming misunderstandings trace back to missing that.

04Delivery guarantees and late data

At least once, at most once, exactly once, and event time versus processing time. These are the topics that separate people who have run a stream from people who have read about one.

Delivery and Time

0/3

At-least-once vs a commit, late event_time, then a watermark that is not the PySpark exercise.

  1. At-least-once14m
  2. Late and out of order14m
  3. Watermarks and Spark14m

By the end of this stage

You can explain what a watermark is for, and what tradeoff you make when you choose how long to wait for late events.

Where people get stuck

Exactly once is widely misunderstood. Learn what it actually guarantees and where the boundary is.

05Streaming and batch together

Real systems are both. The interesting question is how to keep a streaming view and a batch view from disagreeing.

Stream-batch

0/2

Batch is a bounded stream. Replay is seek-and-reread on the same FakeKafkaTopic.

  1. Stream-batch unification12m
  2. Replay from offset12m

By the end of this stage

You can describe how a real architecture serves fresh data and correct historical data at the same time.

06Capstone

Put the pieces together on a realistic event flow with duplicates and late arrivals.

Capstone

0/1

Produce three FakeKafkaTopic records, drop the late one, keep ORD-2 and ORD-3.

  1. Capstone: late messages16m

By the end of this stage

You can design a consumer that is safe to restart and honest about lateness.

How you know it worked

Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.

  • You can explain why a queue sits between producers and consumers.
  • You can describe what an offset is and what happens if a consumer commits before processing.
  • You can explain a watermark and the cost of setting it too tight or too loose.
  • You can argue for batch instead of streaming when it is the right call.

How long it takes

30 minutes a day

about 6 sessions

1 hour a day

about 3 sessions

4 hours a weekend day

about 1 session

Topics here are simulated, so there is no broker to run. Spend your time on offsets and delivery guarantees, because those two modules carry almost all the interview weight and almost all the real-world pain.

These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.

What interviewers are really testing

  • Whether you know ordering is per partition.
  • Whether you can explain duplicate handling with a mechanism, usually idempotent writes or a dedupe key.
  • Whether you understand event time versus processing time, which is the classic senior probe.
  • Whether you can say when streaming is the wrong choice.

Mistakes to avoid on this track

Common mistakes on this track and what to do instead
Common mistakeWhat to do instead
Assuming exactly once means duplicates are impossible everywhere.It is a guarantee about a specific boundary. Your sink still usually needs idempotent writes.
Committing the offset before the work is done.A crash then loses the message silently. Commit after the write, and make the write safe to repeat.
Choosing streaming because it sounds more advanced.Streaming doubles operational burden. If the business can wait an hour, batch is the better engineering decision.

Where to practise this

Production tickets

Duplicate events and offset handling bugs.

System design cases

500M events a day, designed end to end.

Where to go after this

Data Engineering System Design

Streaming decisions are architecture decisions with cost attached.

Data Quality & Observability

Late and duplicate events are a data quality problem as much as a streaming one.

All tracksFull data engineering roadmap45-day plan

Streaming & Message Queues reviews & rating

4.9out of 5
1,240+ student reviews
5 stars
88%
4 stars
9%
3 stars
2%
2 stars
1%
1 star
0%
LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.