Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Streaming & Message Queues

Progress0/13
x

Batch versus Streaming

  • Batch versus streaming12m
  • Streaming architectures12m
  • Exactly-once semantics14m

Why Queues

  • Why a queue12m
  • Topics, partitions, keys14m

Offsets and Groups

  • Offsets and commits14m
  • Consumer groups14m

Delivery and Time

  • At-least-once14m
  • Late and out of order14m
  • Watermarks and Spark14m

Stream-batch

  • Stream-batch unification12m
  • Replay from offset12m

Capstone

  • Capstone: late messages16m
Back to track
  1. Learn
  2. Streaming & Message Queues
  3. Batch versus Streaming
  4. Streaming architectures

Lesson 2 of 13 · Theory first, then run it

Streaming architectures

streamingpythonbeginner12 min

Overview

Lambda runs a batch layer plus a speed layer. Kappa runs only a stream. FakeKafkaTopic is still a Python list. No Kafka broker.

On this page7 sections›
  1. 1The decision
  2. 2What is at stake
  3. 3Option A vs Option B
  4. 4A worked comparison
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The decision

A streaming architecture is the shape of the pipelines that turn events into tables people query. Two names show up in every design review: Lambda and Kappa. They are not products. They are patterns for how many processing paths you maintain.

Lambda keeps two paths. A batch layer recomputes results from stored history (files or warehouse tables). A speed layer processes the latest events so dashboards do not wait for the batch job. A serving layer answers queries, often by combining both.

Kappa keeps one path: a stream on a replayable log. Need yesterday's numbers again? Replay the log through the same code. FakeKafkaTopic in this track is a Python class of lists standing in for that log. No Kafka broker runs in this tab.

What is at stake

Two code paths drift. The night job defines revenue as paid orders. The speed job forgets to exclude cancelled ones. The live chart and the morning report disagree, and the business trusts neither. Lambda's cost is not only extra compute. It is extra definitions to keep in sync.

Kappa's cost is the opposite. You need a log you can replay for days, and stream code that is safe to run twice. If the log is gone after six hours, you cannot rebuild last week. Pick Lambda when you already have strong batch files and only need a speed overlay. Pick Kappa when the log is the system of record you are willing to operate.

Option A vs Option B

Lambda is a kitchen with a slow oven and a microwave: the oven is the accurate batch bake, the microwave is the speed layer for the last few minutes. Kappa is one oven you can run again from the same recipe card (the log).

Lambda (two paths) vs Kappa (one path)
LambdaKappaArchitectureBatch layer: recompute from filesSpeed layer: latest eventsServing layer merges bothReplayable log of eventsOne stream processorServing tables from that stream

Lambda splits history and now. Kappa replays one log through one processor.

PatternLayers you runHow you recomputeUsual risk
LambdaBatch + speed (+ serving)Rerun the batch job on historyTwo definitions of the same metric
KappaStream only (+ serving)Replay the log through the same jobLog retention too short, replay unsafe
Diagram of Lambda Architecture showing Batch Layer, Speed Layer, and Serving Layer
Nathan Marz Lambda Architecture: immutable master events feed the Batch Layer for precomputed views and the Speed Layer for delta streams, unified at the Serving Layer for queries.
Source: Wikimedia CommonsCC BY-SA 4.0

In Python you can sketch the two shapes with functions. These functions do not start a cluster. They name the layers so you remember which work exists.

A worked comparison

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

Python
def lambda_layers():
    return ["batch", "speed"]

def kappa_layers():
    return ["stream"]

result = {
    "lambda": lambda_layers(),
    "kappa": kappa_layers(),
}
print(result)

A tiny speed layer might append events to a running total. A batch layer would re-sum a stored list. Kappa would only keep the append path, and rebuild by walking the list from the start.

Python
# Speed-layer stand-in (latest events). Not a broker.
speed_total = 0.0

def speed_on_event(amount):
    global speed_total
    speed_total += amount
    return speed_total

# Batch-layer stand-in (full recompute from history)
history = [120.50, 40.00, 88.25]

def batch_recompute(amounts):
    return sum(amounts)

print("speed after one click", speed_on_event(12.00))
print("batch from history", batch_recompute(history))
  1. List the questions the serving layer must answer, and how stale the answer may be.
  2. If you already recompute from files every night, a speed overlay is Lambda. Budget time to test both paths against the same grain.
  3. If you can keep a log long enough to rebuild, prefer Kappa: one processor, replay to correct.
  4. In this tab, keep practicing with lists. On a cluster, the log is Kafka (or similar). That homework is still outside this platform.

Common beginner questions

Is Lambda outdated?

Many teams moved toward Kappa because two pipelines were expensive to keep honest. Lambda still appears where a warehouse batch job is already the source of truth and a stream only fills the last hour.

Is Kappa always Kafka?

No. Any durable, replayable log can sit in the middle. Kafka is the common choice in data engineering. Here the log is a Python list.

Where does the serving layer live?

Often a warehouse table, a key-value store, or both. Architecture names the processing paths. Serving is where queries go. Do not skip it in a design sketch.

Two paths means two tests

If you choose Lambda, write the same assertions against batch output and speed output. A speed layer that 'looks live' but disagrees on grain is worse than a slower correct number.

The log is still a list here

FakeKafkaTopic is a Python class of lists. Architecture diagrams transfer. The cluster, retention, and consumer groups do not run in this tab.

What comes next

Either architecture still has to decide what happens when a consumer crashes halfway through an event. The next lesson names the three delivery promises: at-most-once, at-least-once, and exactly-once.

Practice

Run Sample to print the architecture dict. Then complete the exercise: result = {"lambda": ["batch", "speed"], "kappa": ["stream"]} and print it.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Batch versus streamingExactly-once semantics