Unclean leader election allows a broker that is not currently part of the in-sync replica (ISR) set to be elected as the partition leader when all existing ISR members become unavailable. This mechanism restores partition availability for reads and writes, but it causes silent data loss because any committed messages that the out-of-sync broker missed are discarded. Kafka sets unclean.leader.election.enable=false by default to prioritize data durability and consistency over raw uptime.
Why out-of-sync replicas cause data loss
Consider a topic partition where broker 1 is the leader and broker 2 is an in-sync follower. Broker 3 is lagging behind due to heavy disk I/O, so Kafka dropped it from the ISR.
Broker 1 (Leader, ISR): Offsets 0 to 100 Broker 2 (Follower, ISR): Offsets 0 to 100 Broker 3 (Lagging, non-ISR): Offsets 0 to 75
A clear sequence of events explains what happens during a total outage:
- Brokers 1 and 2 experience a sudden power outage on their rack.
- With unclean leader election disabled, the partition becomes offline. Producers and consumers receive errors until either broker 1 or broker 2 recovers, preserving every committed message up to offset 100.
- With unclean leader election enabled, the cluster elects broker 3 as the new leader to keep the partition online. Broker 3 starts accepting new writes starting at offset 76.
Log truncation and consumer inconsistency
When broker 1 eventually reboots and rejoins the cluster, it recognizes broker 3 as the authoritative leader. Broker 1 truncates its local log back to offset 75, permanently deleting committed offsets 76 through 100.
Any downstream consumer that previously read offsets 76 to 100 has processed ghost events that no longer exist in Kafka. Enabling unclean leader election is only acceptable for non-critical metrics or high-volume sensor streams where missing hours of data is preferable to blocking ingestion pipelines.