Rebalance storms occur when consumer group instances repeatedly drop out and rejoin, triggering continuous partition reassignments that halt message consumption across the entire group. The most common trigger is record processing taking longer than max.poll.interval.ms, followed by JVM stop-the-world garbage collection pauses exceeding session.timeout.ms, and rolling container deployments. Preventing rebalance storms requires tuning batch sizes and poll intervals, configuring static group membership, adopting cooperative sticky assignors, or upgrading to the KIP-848 consumer protocol.
Primary triggers of rebalance loops
Three distinct issues typically start a rebalance storm:
- Slow batch processing: If a consumer polls 500 records and spends 6 minutes writing them to a slow database, but max.poll.interval.ms is set to 5 minutes, the coordinator assumes the consumer thread is dead. It revokes its partitions and triggers a rebalance. The evicted consumer finishes its work and calls poll(), forcing yet another rebalance.
- JVM garbage collection pauses: A heavy GC pause freezing the JVM for longer than session.timeout.ms halts the background heartbeat thread, causing the broker to evict the node.
- Rolling restarts: Deploying new application containers using eager rebalancing revokes all partitions from all workers on every pod restart, repeatedly stalling ingestion.
Strategies to eliminate rebalances
Tune processing limits: Lower max.poll.records from 500 to 50 so consumers finish each batch well within max.poll.interval.ms, or increase the interval to accommodate occasional database latency spikes.
Enable static membership: Assign a unique, persistent identifier to each pod via group.instance.id. If a Kubernetes pod restarts within session.timeout.ms, the coordinator preserves its partition assignment rather than triggering a rebalance.
Use cooperative rebalancing: Switch the consumer partition assignor to CooperativeStickyAssignor. Unlike legacy eager assignors that revoke all partitions globally during membership changes, cooperative rebalancing only moves reallocated partitions, allowing unaffected consumers to continue reading without interruption.
Adopt KIP-848: Upgrading to Kafka 4.0 uses the broker-managed consumer protocol, which calculates assignments incrementally on the broker and eliminates stop-the-world rebalance pauses entirely.