Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Data Engineering System Design

Progress0/9
x

How to design

  • Read the brief16m
  • Batch vs stream16m

500M events/day

  • Ingest at 500M events/day18m
  • Storage at 500M events/day18m
  • Failure at 500M events/day18m

Serving and cost

  • Serving analytics16m
  • Cost and FinOps16m

Interview

  • Talk a design in 45 minutes18m

Capstone

  • Capstone: full pipeline design22m
Back to track
  1. Learn
  2. Data Engineering System Design
  3. How to design
  4. Read the brief

Lesson 1 of 9 · Case study

Read the brief

designbeginner16 min

Overview

Volume, SLA, late data, and who reads gold decide the architecture. Name all four before you draw a box.

On this page7 sections›
  1. 1Goal
  2. 2Why this order
  3. 3The checklist
  4. 4Worked pass
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

Goal

A pipeline design brief names four requirements before any architecture diagram: data volume, output freshness (SLA), how late source data can arrive, and which consumers depend on the result.

LakeBench does not run the system you sketch. Ingest, grain, retries, and partitions were practiced on other tracks. This track is the whiteboard where those pieces assemble.

A design brief is the document (or whiteboard section) where you write down the numbers that constrain the system. Without it, you are guessing at tools. Volume tells you how big the pipes need to be. SLA tells you how fast they need to run. Lateness tells you how forgiving the system must be. Consumers tell you what shape the output must take.

Why this order

Without these four numbers, teams pick tools by fashion. A streaming stack for a nightly file, or a batch job for a live counter, wastes money and misses SLAs.

Interviewers and stakeholders both expect you to clarify requirements before drawing boxes. Skipping the brief produces diagrams that look complete but cannot be built or operated.

Imagine you work at a food delivery startup with 500 million daily order events. The growth team wants a real-time order tracker. Finance wants a reconciled revenue report by 08:00. If you build one pipeline for both without separating their requirements, you either over-engineer the finance path (expensive streaming for a nightly report) or under-engineer the tracker (batch lag on a live widget). The brief prevents both mistakes.

The checklist

Four facts before any box
volumeSLAlate dataconsumers

Volume, SLA, lateness, consumers. Draw after these, not instead of them.

Volume is a rate, not a slogan. Five hundred million events per day is about 5,800 per second if arrival is even. Peak hour is often 3 to 5 times that. Ask for events, bytes, distinct keys, and burst shape.

Write the numbers on the board. Guess if needed, and say you guessed.

AskWeak answerUsable answer
How many?A lot of events500M events/day, ~2 KB each, ~1 TB raw
How bursty?Pretty spiky5x peak 18:00-22:00 local
What is an event?User activityclick / view / purchase JSON
Who is the source?The appcheckout service + pixel + 3 vendors

SLA is a clock on a named table. Gold.fct_purchases complete by 08:00 for yesterday is not the same as a click visible within 30 seconds. Quality sla-sli is the vocabulary: SLA is the promise, SLI is the measured lag.

Two clocks that are not the same
eventingestgold readydashboard

Nightly gold by 08:00 versus a 30-second live counter. Name which consumer owns which clock.

Late data is a policy, not an accident. Clicks can arrive after the window you thought was closed. Streaming watermarks and PySpark structured streaming bound lateness in a stream. A nightly batch picks up yesterday's late file on the next run, or you backfill.

  • If late means a file at 03:00 for yesterday, you likely want a night job, not a stream.
  • If late means event_time 02:00 and ingest_time 08:00, you need a watermark and a correction story.
  • If late means the vendor restates last month, you need idempotent loads and a backfill plan.

Consumers decide grain and tolerance. Finance wants additive GMV that reconciles to payments. Growth wants a funnel that can wobble. One gold table cannot serve both if you hide the grain.

Two readers, two requirements
BooksProductwho depends on goldFinance: booksAdditive GMVCorrect by 08:00Replay must not doubleGrowth: funnelCounts can wobbleMinutes, not booksExperiment flags

Same bronze. Different gold. Do not merge them to save a box.

Worked pass

A food delivery startup receives 500 million order events per day from checkout, a web pixel, and a vendor file drop. Before drawing Kafka or BigQuery boxes, capture the four requirements in a dict you can read aloud in an interview.

Four numbers on the board before any product logo.

RequirementValue for this startup
Volume500M events/day, ~2 KB each, 4x peak at dinner
SLAgold.fct_purchases ready by 08:00 UTC
LatenessUp to 6 hours (GPS reconnects late)
ConsumersFinance (exact GMV), Growth (live funnel)
PythonA design brief captured as structured data
# Design brief as a Python dict (conceptual template)
design_brief = {
    "volume": {
        "events_per_day": 500_000_000,
        "avg_event_bytes": 2048,
        "peak_multiplier": 4,
        "sources": ["checkout_service", "pixel", "vendor_hourly"],
    },
    "sla": {
        "table": "gold.fct_purchases",
        "freshness": "yesterday complete by 08:00 UTC",
        "on_miss_page": "data-eng-oncall",
    },
    "lateness": {
        "bound_hours": 6,
        "correction": "next nightly batch picks up late arrivals",
    },
    "consumers": [
        {"team": "finance", "grain": "order_id", "tolerance": "exact reconcile"},
        {"team": "growth", "grain": "session_id", "tolerance": "approx funnel"},
    ],
}
  1. Volume: events/day, bytes, peak multiplier, source count.
  2. SLA: named table, freshness, who pages when it misses.
  3. Lateness: bound in minutes or hours, and the correction path.
  4. Consumers: grain each one needs, and whether they share gold.

A queue is not a personality

Streaming solves a freshness and fan-in problem. Name the four requirements first. The next lesson picks batch versus stream using those numbers.

Do not guess volume without labeling it a guess

Interviewers expect you to state a number, not dodge the question. Say the number, then say whether it is measured or estimated. Wrong with a label is better than vague.

Grain from the SQL track

You practiced grain and additive facts in the SQL track. The consumer section of the brief is asking the same question: what is the primary key of the gold table each reader expects?

PythonBack-of-envelope arithmetic for a design interview
# Quick back-of-envelope rate calculation
events_per_day = 500_000_000
seconds_per_day = 86_400
avg_rate = events_per_day / seconds_per_day  # ~5,787 events/sec

peak_multiplier = 4
peak_rate = avg_rate * peak_multiplier  # ~23,148 events/sec

event_size_kb = 2
peak_throughput_mb = peak_rate * event_size_kb / 1024  # ~45 MB/s inbound at peak

print(f"Average: {avg_rate:.0f} events/sec")
print(f"Peak: {peak_rate:.0f} events/sec")
print(f"Peak throughput: {peak_throughput_mb:.1f} MB/s")

Common beginner questions

How do I know what numbers to write if nobody told me?

Ask. In an interview, state your assumptions out loud and label them guesses. At work, ask the product manager or check monitoring dashboards. The point of the brief is to force the question, not to already know the answer.

Do I need to memorize all these product names?

No. The brief is about the four constraints (volume, SLA, lateness, consumers), not about specific vendor tools. The constraints stay the same whether you use Kafka, Pub/Sub, or Kinesis.

What if the requirements change after I start building?

They will. The brief is a snapshot. Revisit it when volume doubles, new consumers appear, or the SLA tightens. A documented brief makes the change visible instead of silent.

What comes next

The next lesson uses these four numbers to choose between batch and stream processing. Every design decision in this track traces back to the brief you write here.

Practice

On a blank page, write a design brief for a food delivery app with 500 million events per day. Fill in volume, SLA, lateness, and consumers using the template in Worked pass. Run the rate calculation and confirm peak throughput is about 45 MB/s at 4x peak.

No editor on this lesson. The diagrams are the work. Mark it read when you can explain the design out loud.

Rate:
Was this useful?
Data Engineering System DesignBatch vs stream