Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Design a pipeline for IoT sensor data

Pipelines & scenarios · System Design Questions

Design a pipeline for IoT sensor data

Hardpipelines-35
scenarioiottime-seriesstreamingingestion

Question

Design a pipeline for 1 million IoT devices sending readings every 10 seconds.

Solution

First do the arithmetic. One million devices, one reading every 10 seconds, is 100,000 events per second. With 200-byte messages, that is about 20 MB per second, or around 1.7 TB per day raw. That decides most of the design: it is a high-volume, small-message streaming problem.

Devices -> MQTT broker / IoT Core -> Kafka or Pub/Sub
   -> stream processing (validate, enrich, rollup, alert)
   -> hot store (time-series DB)   -> cold store (lake, Parquet)

Ingestion

Devices speak a lightweight protocol, usually MQTT, to a managed broker (AWS IoT Core, Azure IoT Hub, or Google's partner options). Use compact payloads (protobuf or Avro, not verbose JSON), and let devices batch a few readings per message if the battery and latency allow. The broker hands messages to Kafka or Pub/Sub, keyed by device_id, so one device's readings stay ordered. Plan partitions for the peak, not the average.

Processing

A stream job (Flink, Spark Structured Streaming) validates and normalises units, drops duplicates, and computes rollups (per-device 1-minute min, max and average). It also runs alert rules, such as a temperature above a limit for five minutes, and writes alerts to a separate topic.

Devices have bad clocks, and they go offline and send a backlog when they reconnect. So use event time with a generous watermark, and accept that late data will correct older windows. Record both device time and receive time.

Storage in tiers

  • Hot: a time-series store (Bigtable, Cassandra, TimescaleDB, InfluxDB) for "latest value" and the last days of raw data, queried by device and time range.
  • Warm: rollups (1-minute, then hourly) kept for months, which are far smaller than raw.
  • Cold: raw readings as Parquet in object storage, partitioned by date, for analytics and model training. Apply lifecycle rules that move old data to cheaper storage classes.

Operational concerns

Watch ingestion lag and the number of devices reporting. A sudden drop in active devices is an incident signal. Handle a fleet-wide reconnect storm, where a network outage ends and everyone sends at once, with buffers and rate limits. Secure devices with per-device certificates and rotate them.

Retention

Raw for 30 to 90 days in the hot store, rollups for years, and decide this early because it drives cost more than anything else.

PreviousNext