Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Streaming & Message Queues

Progress0/13
x

Batch versus Streaming

  • Batch versus streaming12m
  • Streaming architectures12m
  • Exactly-once semantics14m

Why Queues

  • Why a queue12m
  • Topics, partitions, keys14m

Offsets and Groups

  • Offsets and commits14m
  • Consumer groups14m

Delivery and Time

  • At-least-once14m
  • Late and out of order14m
  • Watermarks and Spark14m

Stream-batch

  • Stream-batch unification12m
  • Replay from offset12m

Capstone

  • Capstone: late messages16m
Back to track
  1. Learn
  2. Streaming & Message Queues
  3. Batch versus Streaming
  4. Batch versus streaming

Lesson 1 of 13 · Theory first, then run it

Batch versus streaming

streamingpythonbeginner12 min

Overview

Nightly warehouse loads are batch. Click-by-click is streaming. Latency, cost, and complexity decide. FakeKafkaTopic in later lessons is a Python class of lists. No Kafka broker runs here.

On this page7 sections›
  1. 1The decision
  2. 2What is at stake
  3. 3Option A vs Option B
  4. 4A worked comparison
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The decision

Batch processing waits until a pile of data exists, then processes the pile in one job. A nightly warehouse load that reads yesterday's orders is batch. The job starts after the day ends. The dashboard updates in the morning.

Streaming processes each event soon after it happens. A click, a payment, a GPS ping: the pipeline does not wait for midnight. Micro-batch sits between the two: small piles on a short timer, often every minute, using many of the same tools as batch.

This track will simulate queues with FakeKafkaTopic, a Python class of lists. No Kafka broker runs in this tab. This first lesson does not need a queue yet. You only need the three names: batch, micro-batch, and streaming, and a rule for when each is enough.

What is at stake

An online store can close the books once a day. Finance does not need each paid order in the warehouse within a second. A fraud check on a card swipe does. If you build a streaming platform for the finance close, you pay for Kafka, consumers, and on-call, to answer a question that a 2 a.m. job already answers.

The other mistake is using only the night job when the business cannot wait. Inventory that updates tomorrow morning will oversell a product that vanished at noon. Latency is a product requirement, not a fashion. Cost and complexity rise as you move from batch toward streaming.

Option A vs Option B

A laundry basket is batch: wait until it is full, then wash. A conveyor at checkout is streaming: each item is scanned as it arrives. A dishwasher that runs every twenty minutes is micro-batch.

How soon the data is ready
Nightly batch(hours)Micro-batch(minutes)Streaming(seconds)

Batch waits for a window to close. Micro-batch shortens the window. Streaming handles events as they arrive.

StyleWhen it runsTypical latencyGood fit
BatchOn a schedule, after a window closesHoursDaily finance, nightly rebuilds
Micro-batchOn a short timer (1 to 15 minutes)MinutesNear-live dashboards without a full stream stack
StreamingOn each event, or tiny groups of eventsSecondsFraud, inventory, click pipelines

Here is the same orders data treated two ways. Batch reads the whole list. Streaming appends one event at a time onto a list that stands in for a log.

Input: three payment events from one evening.

order_ideventpaid_at
ORD-1paid22:14
ORD-2paid22:15
ORD-3paid22:18

Trace the clock

Follow what the user sees at 22:16 in batch versus streaming. Predict who sees the third order first.

A worked comparison

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

Python
# Batch: wait, then process the whole pile
orders = [
    {"order_id": "ORD-1", "amount": 120.50},
    {"order_id": "ORD-2", "amount": 40.00},
    {"order_id": "ORD-3", "amount": 88.25},
]
nightly_revenue = sum(row["amount"] for row in orders)
print("batch revenue", nightly_revenue)

result = ["batch", "micro-batch", "streaming"]
print(result)

A stream does not wait for orders to become a file. Each payment is appended when it happens. Later lessons use FakeKafkaTopic for that log. The object is still a Python list (or list of lists). There is no broker.

Python
# Streaming stand-in: one event at a time onto a list (not a Kafka cluster)
log = []

def on_payment(order_id, amount):
    log.append({"order_id": order_id, "amount": amount})
    return sum(row["amount"] for row in log)

print("after ORD-1", on_payment("ORD-1", 120.50))
print("after ORD-2", on_payment("ORD-2", 40.00))
print("running total", on_payment("ORD-3", 88.25))
  1. Write down the question and the freshest answer the business actually needs.
  2. If tomorrow morning is fine, start with batch. It is simpler to test and cheaper to run.
  3. If a few minutes is fine, consider micro-batch before you introduce a log and consumers.
  4. If seconds matter, plan for a stream: a log, a consumer, and a story for failure and replay.

Common beginner questions

Is micro-batch just streaming with a delay?

Operationally it often uses batch engines on a short schedule. You still have a window, not a per-event consumer. Call it streaming in a meeting and you will confuse the on-call rotation. Keep the names separate.

Why not stream everything?

Streaming systems need a log, consumers, lag monitoring, and a plan for poison messages. That is more moving parts than a scheduled SQL job. Use them when waiting actually costs money or trust.

Does batch mean the data is worse?

No. Batch jobs can be more accurate because they see the whole window and can recompute. Streaming often approximates the latest minute and corrects later. Accuracy and freshness are a tradeoff, not a ranking.

Do not confuse 'real-time' with 'important'

Executives say real-time when they mean 'I do not want to wait until Monday.' Ask for a number: seconds, minutes, or hours. Then pick batch, micro-batch, or streaming.

Before you start

This track assumes you completed Core Python. FakeKafkaTopic is ordinary Python: classes, lists, and dicts. No Kafka broker is installed here.

What comes next

Once you know when a stream is justified, you still have to choose a shape: a batch path plus a speed path (Lambda), or a single replayable stream (Kappa). That is the next lesson.

Practice

Run Sample to print the three styles. Then complete the exercise: store ["batch", "micro-batch", "streaming"] in result and print it.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Streaming & Message QueuesStreaming architectures