Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Hybrid Batch + Streaming Pipeline

HardPro55 min read

Streaming for sub-minute metrics and batch for backfills/heavy aggs, sharing transformation logic (Kappa-inspired).

system-designhybridstreamingbatchkappa

The interview room

You are mid-level. The interviewer wants a hybrid pipeline: streaming for sub-minute metrics and batch for heavy historical work. They also want you to share transformation logic so the two paths do not become two products.

This is not a logo quiz. They want to see that you understand why one path cannot do both jobs well, and how you keep meaning aligned when two engines touch the same business events.

First, the basics

What batch is good at

Batch means: wait until a window of data is complete (a day, an hour), then process it as a set. Strengths: cheap for large scans, easy reprocessing, great for heavy joins and large group-bys, easy to reason about "yesterday's partition is done." Weaknesses: latency. You will not get a sub-minute dashboard from a job that starts at midnight.

What streaming is good at

Streaming means: process events as they arrive (or in small micro-batches). Strengths: freshness, continuous state, alerting, live ops metrics. Weaknesses: heavy historical recompute is awkward, exact month-end finance logic is painful in continuous jobs, and state management adds operational cost.

Why people invented Lambda

Years ago teams wanted both: live dashboards and correct daily numbers. Lambda architecture ran a speed layer (stream) and a batch layer (full recomputes), then merged views. It worked, but often the two codebases diverged. Revenue in the live tile disagreed with finance by Tuesday.

Why Kappa showed up

Kappa says: keep a durable event log, process with one streaming system, and replay the log when you need history. Simpler mental model. Still, some workloads (giant SCD merges, huge warehouse rollups) remain more natural as batch jobs against lake tables.

Modern hybrid (what you should design)

Most companies today do a hybrid that is Kappa-inspired without being dogmatic:

Events --> durable log / Bronze lake
              |                    \
              v                     v
     Stream transforms        Batch transforms
     (shared business rules)  (shared business rules)
              |                     |
              v                     v
     Hot metrics store        Gold daily / monthly marts

Shared rules matter more than shared runtime. Same definition of "completed order," same currency conversion, same refund exclusion. Different engines may still run those rules.

The problem you are designing

Design a hybrid pipeline where:

  • Streaming serves sub-minute product and ops metrics.
  • Batch handles historical backfills and heavy aggregations.
  • Transformation logic is shared (libraries, generated SQL, or one job in two modes).
  • You can explain which path is source of record for which metric class.
  • You can reconcile when stream and batch disagree.

Assume a durable event log (Kafka or equivalent) retained long enough for batch replay into the lake.

What good looks like in the room

You name two audiences. You draw one ingest, two compute paths, two serving stores. You say "stream is approximate / ops; batch is finance SoR" (or whatever policy fits). You show how shared logic is versioned. You describe a reconcile job. You do not pretend one Flink job replaces the warehouse.

Clarifying questions

  • Which metrics truly need sub-minute vs end-of-day?
  • Can batch read the same Kafka topic and/or the same Bronze lake?
  • What is the acceptable error band between live and daily numbers?
  • Who owns metric definitions: analytics eng, DE, or a semantic layer team?
  • How long is the durable log retained?

Constraints and scale

  • Peak event rates may be high, but dashboard QPS is usually modest.
  • Finance cares about correctness over freshness.
  • Product cares about freshness over perfect late-event handling.
  • Backfills of 90 days must not require rewriting stream state from scratch by hand.

Out of scope

  • Multi-region active-active DR design (mention only if asked).
  • Exact vendor bake-off (Flink vs Spark SS) unless they push.
  • Building a full metrics store product from zero.

Why interviewers ask hybrid

They have lived through a company where the live GMV tile and the finance warehouse disagreed for months. They want to hear you prevent that without pretending stream compute is free or that batch can be "realtime enough" by magic.

They also want estimation taste. You should ask which metrics are truly sub-minute. Many "realtime" requests are really "within five minutes" once you push on cost.

A concrete product story

Imagine a marketplace. Ops wants orders-per-minute and payment-error rate on a wallboard. Growth wants funnel conversion with approximate uniques. Finance wants net revenue after refunds for yesterday, locked by 8 AM. Data science wants 180-day training frames rebuilt weekly.

If you design only for ops, finance suffers. If you design only for finance, ops stares at yesterday. Hybrid is the honest architecture for that company.

What "share transformation logic" really means

Sharing does not require one JVM process. It means the rule "refunds reduce net revenue" exists in one place that both paths import, or in tests that fail if either path drifts. Interviewers accept libraries, generated SQL, or dual-mode Spark jobs. They reject "we will be careful" as a strategy.

Good vs weak answers

Weak: "Lambda architecture with batch and speed views" and stop. Good: draw durable log, name SoR per metric class, show reconcile, show versioned shared rules, admit where approximation is labeled in the UI.

Estimation prompts to say out loud

  • Events/sec peak and average
  • Cardinality of metric dimensions (country x product can explode)
  • Hot store write QPS if you flush 1-second buckets
  • Bronze retention days needed for 90-day batch replay
  • Cost of always-on stream vs nightly batch window

Even rough numbers show seniority.

Out of scope reminders you can set

You can park multi-region active-active, exact feature-store product selection, and mobile ingest SDK design unless they ask. Setting scope is part of the interview skill.