Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Real-Time Dashboard with 10-Second Refresh

MediumPro55 min read

Business metrics with 10s refresh using sliding windows, a fast serving store, and WebSocket push.

system-designstreamingdashboardsredis

The interview setup

Design a system for business metrics such as revenue, active users, and conversion with a 10-second refresh SLA. Use sliding-window aggregations, a fast serving store (for example Redis), and WebSocket push so thousands of open dashboards do not poll a warehouse into the ground.

This problem looks friendly. It becomes nasty when someone adds a high-cardinality breakdown "by SKU" without thinking. The interviewer is testing whether you understand preaggregation, bounded metric cardinality, and push versus poll.

Background concepts from first principles

Why polling a warehouse fails

If 2,000 browser tabs each query Snowflake every 10 seconds, you built a distributed denial of service against yourself. Warehouses are great at flexible analytics. They are poor at being a real-time fanout leaf for every open laptop.

Preaggregation

A stream job continuously maintains small metric records: revenue in the current 10-second bucket, approximate unique users, conversion counts. The UI reads those records. Freshness becomes "how often the job flushes" plus "how often clients receive updates."

Sliding versus tumbling windows

Tumbling windows partition time into fixed non-overlapping buckets. Sliding windows hop more often and can overlap. For a 10-second dashboard, tumbling 1-second or 10-second buckets are common; the UI can sum recent buckets.

Serving store

Redis, memory-mapped stores, Druid, Pinot, or ClickHouse can serve hot metrics. Redis is a fine interview default for low-cardinality global metrics. Say when you would graduate to a real-time OLAP store.

Push versus poll

WebSockets (or SSE) push updates from a fanout tier. Polling a lightweight metrics API every 10 seconds can also work at modest scale. Push shines when many clients should stay in sync without stampedes.

The expanded problem

  • Compute sliding or tumbling window metrics for a small set of KPIs
  • Store latest values for fast read
  • Push updates to dashboard clients on a ~10 second cadence
  • Support limited filters (for example by country) without exploding keys
  • Degrade gracefully when the stream lags
Events --> stream windows --> metric keys in Redis
                                 |
                                 v
                           fanout service --WebSocket--> browsers

Constraints and scale prompts

  • Dashboard QPS can be high even when event QPS is moderate, if you let browsers hit the wrong system
  • Metric cardinality is the bomb: global keys are fine; per-SKU keys may not be
  • Exact distinct users at large scale usually means HyperLogLog or batch truth

What good looks like

  • Reject browser-to-warehouse polling immediately
  • Preaggregate into a serving store
  • Bound dimensions
  • Show lag in the UI when consumers fall behind
  • Discuss approximate uniques honestly

Clarifying questions

  • Exact versus approximate unique users?
  • Which filter dimensions are required in v1?
  • Internal-only users or public status pages?
  • Is 10 seconds end-to-end freshness or chart refresh period?
  • How many concurrent dashboard viewers?

Out of scope

  • Pixel-perfect chart library choice
  • Full semantic layer for every BI tool
  • Multi-tenant white-label analytics product

How to use clarifying questions

Ask a few questions that change architecture: SLA definition, source of truth, retention, and tolerance for approximates. If told to decide, state assumptions on the board and proceed.

What good sounds like

Narrate trade-offs while drawing. Name the waiting room for bursts, the transform, the serving surface, and the replay path. Explicitly list out-of-scope items.

Before you draw

Find entry, buffer, transform, serve, and replay. If any is missing, your operational story has a hole.

How to use clarifying questions

Ask a few questions that change architecture: SLA definition, source of truth, retention, and tolerance for approximates. If told to decide, state assumptions on the board and proceed.

What good sounds like

Narrate trade-offs while drawing. Name the waiting room for bursts, the transform, the serving surface, and the replay path. Explicitly list out-of-scope items.

Before you draw

Find entry, buffer, transform, serve, and replay. If any is missing, your operational story has a hole.