Overview
Nightly warehouse loads are batch. Click-by-click is streaming. Latency, cost, and complexity decide. FakeKafkaTopic in later lessons is a Python class of lists. No Kafka broker runs here.
On this page7 sections
The decision
Batch processing waits until a pile of data exists, then processes the pile in one job. A nightly warehouse load that reads yesterday's orders is batch. The job starts after the day ends. The dashboard updates in the morning.
Streaming processes each event soon after it happens. A click, a payment, a GPS ping: the pipeline does not wait for midnight. Micro-batch sits between the two: small piles on a short timer, often every minute, using many of the same tools as batch.
This track will simulate queues with FakeKafkaTopic, a Python class of lists. No Kafka broker runs in this tab. This first lesson does not need a queue yet. You only need the three names: batch, micro-batch, and streaming, and a rule for when each is enough.
What is at stake
An online store can close the books once a day. Finance does not need each paid order in the warehouse within a second. A fraud check on a card swipe does. If you build a streaming platform for the finance close, you pay for Kafka, consumers, and on-call, to answer a question that a 2 a.m. job already answers.
The other mistake is using only the night job when the business cannot wait. Inventory that updates tomorrow morning will oversell a product that vanished at noon. Latency is a product requirement, not a fashion. Cost and complexity rise as you move from batch toward streaming.
Option A vs Option B
A laundry basket is batch: wait until it is full, then wash. A conveyor at checkout is streaming: each item is scanned as it arrives. A dishwasher that runs every twenty minutes is micro-batch.
Batch waits for a window to close. Micro-batch shortens the window. Streaming handles events as they arrive.
| Style | When it runs | Typical latency | Good fit |
|---|---|---|---|
| Batch | On a schedule, after a window closes | Hours | Daily finance, nightly rebuilds |
| Micro-batch | On a short timer (1 to 15 minutes) | Minutes | Near-live dashboards without a full stream stack |
| Streaming | On each event, or tiny groups of events | Seconds | Fraud, inventory, click pipelines |
Here is the same orders data treated two ways. Batch reads the whole list. Streaming appends one event at a time onto a list that stands in for a log.
Input: three payment events from one evening.
| order_id | event | paid_at |
|---|---|---|
| ORD-1 | paid | 22:14 |
| ORD-2 | paid | 22:15 |
| ORD-3 | paid | 22:18 |
Trace the clock
Follow what the user sees at 22:16 in batch versus streaming. Predict who sees the third order first.
A worked comparison
Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.
# Batch: wait, then process the whole pile
orders = [
{"order_id": "ORD-1", "amount": 120.50},
{"order_id": "ORD-2", "amount": 40.00},
{"order_id": "ORD-3", "amount": 88.25},
]
nightly_revenue = sum(row["amount"] for row in orders)
print("batch revenue", nightly_revenue)
result = ["batch", "micro-batch", "streaming"]
print(result)A stream does not wait for orders to become a file. Each payment is appended when it happens. Later lessons use FakeKafkaTopic for that log. The object is still a Python list (or list of lists). There is no broker.
# Streaming stand-in: one event at a time onto a list (not a Kafka cluster)
log = []
def on_payment(order_id, amount):
log.append({"order_id": order_id, "amount": amount})
return sum(row["amount"] for row in log)
print("after ORD-1", on_payment("ORD-1", 120.50))
print("after ORD-2", on_payment("ORD-2", 40.00))
print("running total", on_payment("ORD-3", 88.25))- Write down the question and the freshest answer the business actually needs.
- If tomorrow morning is fine, start with batch. It is simpler to test and cheaper to run.
- If a few minutes is fine, consider micro-batch before you introduce a log and consumers.
- If seconds matter, plan for a stream: a log, a consumer, and a story for failure and replay.
Common beginner questions
Is micro-batch just streaming with a delay?
Operationally it often uses batch engines on a short schedule. You still have a window, not a per-event consumer. Call it streaming in a meeting and you will confuse the on-call rotation. Keep the names separate.
Why not stream everything?
Streaming systems need a log, consumers, lag monitoring, and a plan for poison messages. That is more moving parts than a scheduled SQL job. Use them when waiting actually costs money or trust.
Does batch mean the data is worse?
No. Batch jobs can be more accurate because they see the whole window and can recompute. Streaming often approximates the latest minute and corrects later. Accuracy and freshness are a tradeoff, not a ranking.
Do not confuse 'real-time' with 'important'
Executives say real-time when they mean 'I do not want to wait until Monday.' Ask for a number: seconds, minutes, or hours. Then pick batch, micro-batch, or streaming.
Before you start
This track assumes you completed Core Python. FakeKafkaTopic is ordinary Python: classes, lists, and dicts. No Kafka broker is installed here.
What comes next
Once you know when a stream is justified, you still have to choose a shape: a batch path plus a speed path (Lambda), or a single replayable stream (Kappa). That is the next lesson.
Practice
Run Sample to print the three styles. Then complete the exercise: store ["batch", "micro-batch", "streaming"] in result and print it.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.