Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Sessionization System for User Activity

HardPro55 min read

Group user events into sessions with a 30-minute inactivity gap; handle out-of-order events, late data, and large keyed state.

system-designstreamingsessionizationevent-time

The interview setup

Design a streaming system that groups user events into sessions using a 30-minute inactivity gap. You must handle out-of-order events, late arrivals, and state for millions of concurrent users.

This is one of the classic stream processing interviews. Candidates who quietly sessionize on processing time invent fiction. Candidates who never say "watermark" get a follow-up that hurts. Candidates who ignore state size design a demo, not a system.

What the interviewer wants:

  • A crisp event-time definition of a session
  • A watermark and lateness policy
  • A plan for keyed state and TTL
  • A late-data story after sessions close
  • Deterministic identifiers so sinks can be idempotent

Background concepts from first principles

What is a session?

A session is a burst of activity by one user separated by idle time. If the gap between consecutive events in event time exceeds thirty minutes, start a new session.

Example: events at 1:00, 1:10, and 1:35 stay one session because the gaps are under thirty minutes. An event at 2:10 starts a new session.

Product uses sessions for funnel analysis, engagement time, attribution, and personalization. Wrong splits corrupt all of that.

Event time versus processing time

Event time is when the action happened according to the producer clock (or a trusted server timestamp assigned at the edge). Processing time is when your job saw the event.

Mobile phones go offline. Partitions lag. Retries reorder. If you close sessions with wall-clock now, you are measuring your pipeline health, not user behavior.

Watermarks

A watermark is the engine's working belief about event-time progress: "I think I have seen most data up to time T." Windows and sessions use that belief to decide when they can close.

Watermarks are heuristics. They can be wrong. Your design should admit that.

Allowed lateness

Allowed lateness keeps session state around after the watermark passes so late events can still join for a bounded time. Too large and state explodes. Too small and you split sessions incorrectly or drop updates.

Keyed state

Open sessions live in state keyed by user_id. Millions of open sessions means RocksDB-sized state, checkpoint costs, and a TTL story after close.

The expanded problem

Build a pipeline that:

  • Assigns session ids using a thirty-minute inactivity gap in event time
  • Handles out-of-order events within a stated policy
  • Emits session aggregates when sessions close
  • Scales to millions of concurrent open sessions
  • Supports reprocessing after bugs
User events (event_time)
        |
        v
      Kafka (key=user_id)
        |
        v
 Sessionizer (event-time gap = 30m)
        |
        +--> session facts sink
        +--> late side output

Constraints and scale prompts

  • State is the scarce resource, not raw bytes
  • Checkpoint frequency versus recovery time is a real trade-off
  • Mobile clock skew can poison event_time if unchecked
  • Downstream BI often wants stable session grains

What good looks like

  • Define the session in one sentence with event time
  • Draw watermark plus lateness
  • Estimate state for millions of keys
  • Explain what happens to very late events
  • Mention deterministic session ids and replay

Clarifying questions

  • Is event_time from the client clock or assigned at the ingest server?
  • How late is acceptable: minutes, hours, or a day?
  • Should late events revise closed sessions or only land in a corrections path?
  • Which aggregates are required on close: duration, event count, revenue?
  • Do we need exact replay for a bad day?

Out of scope

  • Full cross-device identity graph stitching (mention as a sequel)
  • Designing the mobile SDK in detail
  • Pixel-perfect dashboard layout
  • Exactly-once theology without a sink discussion (keep it practical)

How to use your clarifying questions

Ask two or three high-value questions early, not twelve. Prioritize questions that change the diagram: latency SLA definition, source of truth, retention, and whether approximate answers are allowed. Park niche questions until the deep dive.

If the interviewer refuses to answer and says "you decide," state your assumption explicitly and proceed. Write the assumption on the board so it stays visible.