In the classic consumer group protocol, the broker acting as group coordinator manages group membership and heartbeats, but delegates the partition assignment calculation to the designated group leader consumer. The newer KIP-848 protocol moves partition assignment computation entirely onto the broker coordinator, eliminating heavy stop-the-world rebalance pauses. Health is maintained via consumer background heartbeats governed by session.timeout.ms, while record processing deadlocks are caught by max.poll.interval.ms.
The classic partition assignment cycle
When consumers join a group, they send JoinGroup requests to the coordinator broker. The coordinator elects the first consumer to respond as the group leader and returns the list of all active members to that leader. The group leader consumer executes a partition assignor strategy:
- RangeAssignor: Assigns contiguous partition ranges per topic, which can cause consumer imbalance across multiple topics.
- RoundRobinAssignor: Distributes partitions evenly across all consumers.
- CooperativeStickyAssignor: Retains existing partition assignments and reallocates partitions incrementally without stopping unaffected consumers.
The leader sends the resulting assignment plan back to the coordinator in a SyncGroup request, and the coordinator distributes the individual assignments to each member.
Heartbeats and processing timeouts
Consumers run a dedicated background heartbeat thread that pings the coordinator. If no heartbeat arrives within session.timeout.ms (default 45 seconds), the coordinator considers the instance dead and kicks it out of the group.
Separately, the application main thread must call poll() regularly. If processing a batch of records against a slow database takes longer than max.poll.interval.ms (default 5 minutes), the consumer voluntarily abandons its partitions, triggering a rebalance.
Broker-side assignment in KIP-848
The next-generation protocol introduced in KIP-848 removes client-side assignors. Consumers simply report their state to the broker, and the group coordinator computes incremental partition assignments directly, avoiding cluster-wide pauses during scaling events.