A rebalance is when Kafka redistributes partition assignments among members of a consumer group.
When it happens
- A consumer joins the group
- A consumer leaves (graceful close or crash)
- A consumer misses heartbeats / session timeout
- Subscription changes (subscribe to more/fewer topics)
- Partitions added to a topic (metadata change)
Before: C1→[P0,P1] C2→[P2,P3] C3 joins After: C1→[P0] C2→[P1,P2] C3→[P3] (example assignment)
Why it hurts
During classic rebalances, consumers may stop processing assigned partitions temporarily. Frequent rebalances → lag spikes and "rebalance storms."
What consumers should do
- On revoke: commit offsets, flush local state carefully.
- On assign: seek/initialize for new partitions.
- Keep processing time per poll reasonable; call
polloften enough for heartbeats (or use cooperative patterns / separate heartbeat thread in modern clients).
Ops tips to reduce churn
- Stable
group.idand deployment rollout strategy - Sensible
session.timeout.ms/max.poll.interval.ms - Avoid unnecessary subscribe changes
- Consider static membership and cooperative sticky assignors
Interview tip: "Rebalance = remap partitions to group members. Triggered by membership or subscription change. Too many rebalances hurt lag."