The interview setup
Design a pipeline for 10TB of daily clickstream. Use Protobuf serialization, two-stage processing (Flink sessionization + Spark daily aggregations), and cost optimization with spot instances.
The interviewer is testing whether you understand why one mega-job is a bad idea, why binary schemas matter at this volume, and how spot compute changes failure design.
Background concepts from first principles
Why clickstream gets expensive
JSON click events are easy to debug and expensive to store and parse at multi-TB/day. CPU burns on parsing. Object storage burns on fat bytes. Downstream jobs burn on reading them again.
Protobuf (and friends)
Protobuf is a compact binary schema with evolving field rules. You get smaller payloads and stronger contracts than free-form JSON. Avro and similar formats play the same role. The interview point is schema discipline + compact encoding, not brand loyalty.
Two stages: stateful vs embarrassingly parallel
Sessionization needs event-time state per user. That is a streaming or micro-batch stateful job.
Daily aggregations (revenue by country, top pages) are often group-bys over a date partition. Spark batch on spot fits well.
Mixing both in one untunable monster job makes ownership, retries, and cost controls worse.
Spot / preemptible compute
Spot machines are cheap and can disappear. Fine when checkpoints and atomic outputs exist. Painful for long state rebuilds if you did not plan for interruption.
The expanded problem
- Ingest and store daily clickstream efficiently (Protobuf into columnar lake files)
- Sessionize events with Flink (or equivalent)
- Compute daily aggregations in Spark
- Use spot where safe
- Support reprocessing a day
Raw Protobuf landing (S3)
|
v
Flink sessionization (stateful) --> session tables
|
v
Spark daily aggs (spot) --> gold martsScale prompts
- 10TB/day raw ≈ ~116 MB/s average; peaks higher
- Compression to Parquet/ZSTD can cut storage sharply after decode
- Session state sizing depends on concurrent users, not only TB/day
What good looks like
- Cost of JSON called out early
- Clear stage boundary between sessions and daily aggs
- Spot strategy with checkpoint/commit story
- Reprocess-by-date plan
- Schema evolution awareness
Clarifying questions
- Session definition and allowed lateness?
- Retention of raw versus sessions?
- Cloud spot characteristics and interruption rate?
- SLA for morning aggregates?
- Peak versus average traffic shape?
Out of scope
- Building a new serialization format
- Perfect ML on clickstream
- Multi-cloud networking deep dive
Clarifying questions strategy
Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.
What good looks like verbally
Narrate why stages exist. Name atomic publish. Name how reruns stay safe. Mention what is out of scope.
Scope control
Batch interviews reward candidates who protect the critical morning path and park nice-to-haves. Say what can wait until after the SLA table lands.
Clarifying questions strategy
Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.
What good looks like verbally
Narrate why stages exist. Name atomic publish. Name how reruns stay safe. Mention what is out of scope.
Scope control
Batch interviews reward candidates who protect the critical morning path and park nice-to-haves. Say what can wait until after the SLA table lands.
Clarifying questions strategy
Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.
What good looks like verbally
Narrate why stages exist. Name atomic publish. Name how reruns stay safe. Mention what is out of scope.
Scope control
Batch interviews reward candidates who protect the critical morning path and park nice-to-haves. Say what can wait until after the SLA table lands.
Clarifying questions strategy
Ask questions that change grain, retention, SLA, or cost. If told to decide, state assumptions explicitly and continue.