Monitoring an Apache Kafka cluster requires tracking core broker availability, data replication health, consumer lag, and host hardware resources. The highest-priority alerting metrics are under-replicated partitions, offline partitions, and active controller count, while consumer lag reveals downstream processing bottlenecks. Production teams capture these metrics using the Prometheus JMX Exporter, Burrow, and Cruise Control to maintain cluster health and detect issues before outages occur.
Critical broker and replication metrics
Set up immediate paging alerts on the following broker metrics:
- UnderReplicatedPartitions: The number of partitions where the in-sync replica count is less than the configured replication factor. Any value above zero indicates a dead broker, disk failure, or network partition.
- OfflinePartitions: Partitions without an elected leader. When this metric is greater than zero, affected partitions cannot accept writes or serve reads.
- ActiveControllerCount: Must equal exactly 1 across the entire cluster. A value of 0 or a number greater than 1 signals controller election failure or split-brain metadata states.
- IsrShrinksPerSec and IsrExpandsPerSec: Spikes in ISR shrink rate indicate brokers falling behind due to disk saturation or network latency.
Consumer lag and performance metrics
Monitor consumer group health and broker traffic:
- Consumer lag: The difference between the latest partition log end offset and the consumer committed offset. Steadily rising lag flags stalled consumers or downstream processing bottlenecks.
- TotalTimeMs (p99): Measures total request time for produce and fetch requests, exposing broker disk I/O or network contention.
- Disk space usage: Alert before broker disks reach 80% capacity to prevent emergency shutdowns or log cleaner failures.
- Network throughput and CPU: Kafka is typically network-bound; monitor network interface card saturation on broker hosts.
Production monitoring tooling
Export broker JMX metrics into Prometheus using the Prometheus JMX Exporter and build Grafana alert dashboards. Use specialized tools like Burrow to monitor consumer lag trends without placing administrative polling load on brokers, and deploy Cruise Control to automate partition rebalancing across unevenly loaded brokers.