Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Unclean leader election

Kafka · Internals

Unclean leader election

Hardkafka-37
unclean-leader-electionhigh-availabilitydata-lossfault-tolerance

Question

What is unclean leader election, and why is it disabled by default?

Solution

Unclean leader election allows a broker that is not currently part of the in-sync replica (ISR) set to be elected as the partition leader when all existing ISR members become unavailable. This mechanism restores partition availability for reads and writes, but it causes silent data loss because any committed messages that the out-of-sync broker missed are discarded. Kafka sets unclean.leader.election.enable=false by default to prioritize data durability and consistency over raw uptime.

Why out-of-sync replicas cause data loss

Consider a topic partition where broker 1 is the leader and broker 2 is an in-sync follower. Broker 3 is lagging behind due to heavy disk I/O, so Kafka dropped it from the ISR.

Broker 1 (Leader, ISR):      Offsets 0 to 100
Broker 2 (Follower, ISR):    Offsets 0 to 100
Broker 3 (Lagging, non-ISR): Offsets 0 to 75

A clear sequence of events explains what happens during a total outage:

  • Brokers 1 and 2 experience a sudden power outage on their rack.
  • With unclean leader election disabled, the partition becomes offline. Producers and consumers receive errors until either broker 1 or broker 2 recovers, preserving every committed message up to offset 100.
  • With unclean leader election enabled, the cluster elects broker 3 as the new leader to keep the partition online. Broker 3 starts accepting new writes starting at offset 76.

Log truncation and consumer inconsistency

When broker 1 eventually reboots and rejoins the cluster, it recognizes broker 3 as the authoritative leader. Broker 1 truncates its local log back to offset 75, permanently deleting committed offsets 76 through 100.

Any downstream consumer that previously read offsets 76 to 100 has processed ghost events that no longer exist in Kafka. Enabling unclean leader election is only acceptable for non-critical metrics or high-volume sensor streams where missing hours of data is preferable to blocking ingestion pipelines.

PreviousNext