Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Tiered storage

Kafka · Operations & Scenarios

Tiered storage

Mediumkafka-56
tiered-storagekip-405storage-architecturecloud-storage

Question

What is Kafka tiered storage?

Solution

Kafka Tiered Storage (introduced under KIP-405 and production-ready since Apache Kafka 3.9) separates cluster compute from durable storage by offloading inactive log segments from local broker disks to scalable cloud object storage such as Amazon S3 or Google Cloud Storage. Active and recent log segments remain on fast local SSDs or NVMe drives for low-latency real-time consumption, while cold historical data is archived remotely for months or years. This architecture enables cost-effective long retention and drastically accelerates broker partition rebalancing, though reading cold historical data incurs higher fetch latency.

The two storage tiers

Tiered storage splits partition management into two distinct layers:

  • Local tier: Ingests all producer writes directly to fast local disks. Consumers operating at the head of the log read hot records straight from the Linux OS page cache with sub-millisecond response times.
  • Remote tier: Once an active segment fills up and closes, a background Remote Storage Manager uploads the closed log and index files to an object storage bucket.

Local retention policies can delete local segment files after hours or days, freeing local disk capacity while remote retention policies keep the data searchable in object storage for months.

Faster broker rebalancing and operational recovery

In traditional Kafka clusters, expanding a cluster or replacing a degraded broker requires copying hundreds of gigabytes of partition history over the network from peer brokers. With tiered storage enabled, partition reassignments only copy the small local storage tier. The new broker simply points its metadata to the remote segment objects stored in S3, reducing broker rebalance times from hours down to minutes.

The read latency trade-off

While tiered storage dramatically lowers cloud disk costs, reading historical segments requires brokers to fetch remote chunks over HTTP APIs from object storage. Batch backfills and historical replay jobs experience higher initial fetch latency compared to reading local NVMe drives.

PreviousNext