Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Log Aggregation System (10,000 Servers)

HardPro55 min read

Collect logs from 10k servers at ~250 GB/sec aggregate, support sub-second error search for 7 days, and archive for 1 year.

system-designloggingsearchstorageestimation

Interview framing

You are designing a log aggregation platform for a large fleet. Interview length about 55 minutes. The prompt will throw an extreme ingest number at you on purpose. Strong candidates negotiate the number, then separate hot searchable errors from cheap long-term archive.

The prompt:

Collect logs from 10,000 servers. Each server generates about 50,000 log lines per second in the aggressive version of the prompt, with aggregate volume around 250 GB/sec. You need full-text search on error logs with sub-second latency for the last 7 days, and archive all logs for 1 year.

Background from first principles

Logs are append-only records of what software did: info, warn, error, debug. Operators search them during incidents. Compliance may require long retention. Those two needs fight each other. Full-text indexes are expensive. Object storage is cheap. If you put every line for a year into Elasticsearch, you are designing a second company.

Agents on hosts (Fluent Bit, Vector, and similar) must protect the application. If the logging backend is down, the agent buffers then drops with metrics. Logging must never deadlock checkout.

Understand the access patterns before picking stores:

  • Incident response: recent errors, filter by service and time, sub-second.
  • Postmortem deep dive: maybe weeks ago, can be slower.
  • Compliance audit: prove retention, rarely interactive search.

Those are different systems wearing one product name.

Expanded problem and constraints

  • Fleet: 10k servers.
  • Prompt scale: about 250 GB/sec aggregate. Treat as interview extreme. Say if you negotiate down to a softer 1 to 5 GB/sec while keeping the same shape.
  • Hot path: error (and maybe warn) logs searchable for 7 days, typical queries under about 1 second.
  • Cold path: all levels archived 1 year with compression and lifecycle tiering.
  • Filters: service, host, severity, time range.
  • Agents: restarts should not create silent huge gaps. Detect gaps.
  • Cost: full-text indexing everything for a year is usually not viable. Say so.

What good looks like

You challenge 250 GB/sec politely. You route severity=error to a hot search cluster with index lifecycle management. You ship everything to object storage. You design agent backpressure and per-service quotas. You explain how to find logs from day 60 (cold scan or rehydrate), not by keeping a year of hot indexes.

You also mention structured JSON fields so filters do not depend on brittle free-text alone.

Clarifying questions

  • Is 250 GB/sec real or a stress number we can negotiate?
  • Must all log levels be searchable, or only error/warn?
  • Structured JSON logs vs free text?
  • Retention and compliance (PII or secrets in logs)?
  • Multi-region apps?
  • Who pays if one noisy service emits 80% of volume?
  • Are trace ids standard across services?

Scale and estimation prompts

  • Even at 5 GB/sec, 7 days of raw is huge; indexed copies multiply cost.
  • Errors are a small fraction of total volume. Index selectively.
  • 1 year archive belongs in object storage with gzip/zstd and cold tiers, not forever-hot search nodes.
  • Search QPS during incidents spikes; size hot cluster for query concurrency on the error subset.
  • Show the shape at negotiated scale and at prompt scale so the interviewer sees dials, not panic.

Work a softer estimate out loud: 10k hosts x 1 MB/s is 10 GB/s. Still enormous. Sampling debug and indexing only errors remains mandatory.

Out of scope for v1

  • Perfect fidelity debug search for every line for 365 days.
  • Blocking application threads when the logging sink is slow.
  • A single cluster that is both unlimited ingest and unlimited retention full-text.
  • Guaranteeing zero log loss under total backend outage (prefer drop with loud metrics).
  • Building a full APM traces product unless asked (logs vs traces vs metrics).