Consumer lag = how far behind a consumer group is on a partition.
lag = (high watermark / end offset) - (consumer's current/committed offset)
If the log end is offset 10_000 and the group is at 9_700, lag ≈ 300 messages.
Why it matters
- High lag → stale dashboards, delayed fraud checks, late warehouse loads.
- Growing lag → consumer cannot keep up with produce rate (or is stuck).
- Sudden lag spike → outage, rebalance storm, bad deploy, or produce burst.
Common causes
- Too few consumers / partitions for the volume
- Slow processing (heavy DB writes per message)
- Long GC / resource starvation
- Frequent rebalances interrupting consumption
- Downstream dependency outages causing retries/backpressure
- Hot partition (skewed keys) overloading one consumer
What to check
1. Lag by topic/partition/group (Burrow, Confluent Control Center, Prometheus exporters, cloud console). 2. Is lag on one partition (skew) or all (global slowness)? 3. Consumer CPU, poll time, processing time. 4. Rebalance frequency and errors in logs. 5. Produce rate vs consume rate.
Healthy: lag low and stable Warning: lag rising over time Incident: lag huge + SLO breach
Interview tip: Define lag as end offset minus consumer offset, then list scale-out, fix skew, and speed up processing as remediation levers.