Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Batch ETL for 10TB Daily Clickstream

HardPro55 min read

Process 10TB/day clickstream with Protobuf, Flink sessionization + Spark daily aggregations, and spot-instance cost control.

system-designbatchclickstreamcost-optimizationflinkspark

Warm-up Trace

Warm-up Trace

The interview setup

Design a pipeline for 10TB of daily clickstream. Use Protobuf serialization, two-stage processing (Flink sessionization + Spark daily aggregations), and cost optimization with spot instances.

The interviewer is testing whether you understand why one mega-job is a bad idea, why binary schemas matter at this volume, and how spot compute changes failure design.

Background concepts from first principles

Why clickstream gets expensive

JSON click events are easy to debug and expensive to store and parse at multi-TB/day. CPU burns on parsing. Object storage burns on fat bytes. Downstream jobs burn on reading them again.

Protobuf (and friends)

Protobuf is a compact binary schema with evolving field rules. You get smaller payloads and stronger contracts than free-form JSON. Avro and similar formats play the same role. The interview point is schema discipline + compact encoding, not brand loyalty.

Two stages: stateful vs embarrassingly parallel

Sessionization needs event-time state per user. That is a streaming or micro-batch stateful job.

Daily aggregations (revenue by country, top pages) are often group-bys over a date partition. Spark batch on spot fits well.

Mixing both in one untunable monster job makes ownership, retries, and cost controls worse.

Spot / preemptible compute

Spot machines are cheap and can disappear. Fine when checkpoints and atomic outputs exist. Painful for long state rebuilds if you did not plan for interruption.

The expanded problem

  • Ingest and store daily clickstream efficiently (Protobuf into columnar lake files)
  • Sessionize events with Flink (or equivalent)
  • Compute daily aggregations in Spark
  • Use spot where safe
  • Support reprocessing a day
Raw Protobuf landing (S3)
        |
        v
Flink sessionization (stateful) --> session tables
        |
        v
Spark daily aggs (spot) --> gold marts

Scale prompts

  • 10TB/day raw ≈ ~116 MB/s average; peaks higher
  • Compression to Parquet/ZSTD can cut storage sharply after decode
  • Session state sizing depends on concurrent users, not only TB/day

What good looks like

  • Cost of JSON called out early
  • Clear stage boundary between sessions and daily aggs
  • Spot strategy with checkpoint/commit story
  • Reprocess-by-date plan
  • Schema evolution awareness

Clarifying questions

  • Session definition and allowed lateness?
  • Retention of raw versus sessions?
  • Cloud spot characteristics and interruption rate?
  • SLA for morning aggregates?
  • Peak versus average traffic shape?

Out of scope

  • Building a new serialization format
  • Perfect ML on clickstream
  • Multi-cloud networking deep dive

Clarifying questions strategy

Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.

What good looks like verbally

Narrate why stages exist. Name atomic publish. Name how reruns stay safe. Mention what is out of scope.

Scope control

Batch interviews reward candidates who protect the critical morning path and park nice-to-haves. Say what can wait until after the SLA table lands.

Clarifying questions strategy

Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.

What good looks like verbally

Narrate why stages exist. Name atomic publish. Name how reruns stay safe. Mention what is out of scope.

Scope control

Batch interviews reward candidates who protect the critical morning path and park nice-to-haves. Say what can wait until after the SLA table lands.

Clarifying questions strategy

Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.

What good looks like verbally

Narrate why stages exist. Name atomic publish. Name how reruns stay safe. Mention what is out of scope.

Scope control

Batch interviews reward candidates who protect the critical morning path and park nice-to-haves. Say what can wait until after the SLA table lands.

Clarifying questions strategy

Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.