Interview framing
You are in a mid-level data engineering interview. The interviewer asks you to compare Lambda Architecture and Kappa Architecture for analytics pipelines.
They are not fishing for buzzwords. They want to know whether you have lived the pain of two pipelines that disagree, whether you understand replay, and whether you can recommend something a mid-size team can operate without drowning.
A strong answer starts with the product problem that created these patterns. Then you name the layers. Then you tell the failure story. Then you pick for *this* company.
Background: what existed before Lambda
Before "Lambda Architecture" was a phrase, most analytics looked like this.
Events and logs landed in files or a database dump. Overnight, a batch job scanned yesterday's data and wrote daily aggregates into a warehouse. Product managers opened dashboards at 8 AM and saw yesterday's truth.
That model was correct for many questions. It was also late. Product started asking for "active users in the last five minutes," "error rate right now," and "did the checkout funnel break after the deploy." Waiting until tomorrow was not acceptable.
Teams bolted on a second path: a streaming job that approximated the same metrics in seconds. The warehouse stayed the "official" numbers. The stream fed live dashboards. For a while everyone was happy.
Then the numbers diverged. Batch said 1.02M daily actives. Stream said 980K. Finance trusted batch. Ops trusted stream. Nobody could explain the gap in a sentence. That disagreement is the historical problem Lambda tried to organize, and it is still the scoring center of this interview question.
Background: first principles before the names
You need three ideas before you draw boxes.
1. Correctness vs freshness
Some questions need exact, reproducible answers over large history (revenue by day with late refunds applied). Some questions need a fresh approximate answer (errors in the last 60 seconds). One engine and one schedule rarely optimize both.
2. Two implementations of the same business rule
If you compute "active user" in SQL with one set of window rules, and again in a streaming job with different event-time semantics, you have two products pretending to be one metric. Drift is not a bug in Kafka. Drift is a process failure.
3. Replay as a first-class operation
When you discover a bug in how you counted, can you rebuild last week's metrics from durable raw inputs? If yes, you can fix history. If no, you live with wrong dashboards forever or with painful one-off backfills.
Lambda and Kappa are two different answers to how you split correctness, freshness, and replay.
Full problem statement
Compare Lambda and Kappa for a mid-size company:
- Tens of thousands of events per second (not hyperscale).
- A few data engineers maintain pipelines.
- Product wants some live metrics and some trusted daily numbers.
- Clickstream is a fine concrete domain: page views, clicks, purchases.
Explain:
- Lambda's batch layer, speed layer, and serving layer.
- Kappa's single streaming pipeline plus replay from a durable log.
- When you recommend each.
- What goes wrong with Lambda in practice, especially code divergence.
- How corrections and backfills work in each.
What "good" looks like in the room
You tell a short story about disagreeing metrics before you define terms. You draw both architectures on the board. You walk a clickstream example through each path. You spend real time on code divergence with a vivid failure. You discuss replay cost. You make a recommendation tied to team size and whether a durable log with weeks of retention is feasible. You do not declare Kappa always better.
Clarifying questions to ask
- Do we already have a mature batch warehouse and SQL culture?
- Is a durable, replayable log (Kafka or similar) already in place, and how long is retention?
- How often do metric definitions change?
- Latency needs: seconds vs minutes vs hours?
- Who is the consumer of "truth" for finance vs ops dashboards?
- Can the streaming engine economically replay 30-90 days of history?
Scale and estimation prompts
Say rough numbers out loud even if the prompt is qualitative.
- 50K events/sec × 500 bytes ≈ 25 MB/sec ≈ ~2 TB/day raw before compression.
- Kafka with 3× replication multiplies network and disk.
- Replaying 30 days at catch-up speed may need far more compute than the live job.
- Serving-layer merge in Lambda adds read complexity: how do apps combine batch truth and speed approximations?
What is out of scope
- Picking a specific vendor logo as the answer (Kafka vs Pulsar, Flink vs Spark) without tying it to constraints.
- Designing a full multi-region DR plan unless asked.
- Deep algorithm design for exact distincts at hyperscale.
- Claiming one architecture is always modern and the other is obsolete.
Keep the conversation on layers, divergence, replay, and a defendable recommendation for a mid-size DE team.