A Kafka partition is stored on the broker filesystem as a dedicated directory containing immutable log segment files paired with offset and timestamp index files. Retention policies evaluate and delete entire closed segments based on time or size limits, but Kafka will never delete the currently active segment that is receiving writes. Because data deletion happens strictly at the segment boundary, messages can stay on disk much longer than retention.ms if segment.ms or segment.bytes has not yet closed the active file.
Structure of a partition directory
Within the broker storage path, each partition has a directory formatted as topic name followed by partition ID, such as orders-0. Inside that directory, data is split into segments named after their base offset:
- The
.logfile stores the raw serialized message bytes. - The
.indexfile stores a memory-mapped sparse index mapping logical offsets to physical byte positions in the log. - The
.timeindexfile stores a sparse index mapping message timestamps to logical offsets for time-based lookups.
The highest-numbered segment is the active segment. All new producer records append directly to this active file.
Why data outlives retention.ms
A background cleaner thread periodically inspects segments against retention.ms (time-based limit) and retention.bytes (size-based limit). However, two strict rules govern this deletion process:
- Retention rules only purge closed, inactive segments. The active segment is never touched, regardless of how old its oldest message is.
- Segments roll over into closed status only when they reach segment.bytes (default 1 GB) or when segment.ms elapses (default 7 days).
If a low-throughput topic receives only 100 MB of data per week and has retention.ms=3 days with default segment.bytes=1 GB and default segment.ms=7 days, the active segment will not roll over until day seven. The messages written on day one will survive on disk for over a week before the segment closes and becomes eligible for cleanup.