Lambda architecture keeps two paths to answer queries on the same data:
1. Batch layer: periodic recomputation over the full master dataset (accurate, slower) 2. Speed layer: streaming approximations for recent data (fast, may be incomplete) 3. Serving layer: merges batch views + real-time views for the user
+--> batch (nightly Spark) ----+
all events -----+ +--> serving merge --> query
+--> speed (stream) -----------+Why it existed
Early streaming engines were weaker; batch was trusted for correctness while stream filled the freshness gap.
Pain
You maintain two codepaths for similar business logic. Bugs diverge. Ops cost doubles.
Interview tip: Define batch + speed + serving, then say many teams moved toward Kappa or unified batch/stream APIs to avoid dual logic.