Kafka producer retries handle transient broker leader re-elections and network timeouts by automatically resending failed record batches within a bounded delivery window. Modern Kafka sets retries to Integer.MAX_VALUE by default, relying on delivery.timeout.ms (default 120,000 milliseconds) as the upper bound for total elapsed time before declaring a permanent failure. Without producer idempotence enabled, retries risk duplicating records or altering event order, while terminal timeouts trigger the user-defined completion callback with an exception.
Time-based retry lifecycle
Rather than counting discrete retry attempts, the producer enforces a comprehensive time window:
- delivery.timeout.ms: The maximum time allocated for a record batch to be acknowledged, beginning the moment send() places it in the accumulator.
- request.timeout.ms (default 30 seconds): The timeout for a single network attempt to a broker.
- retry.backoff.ms (default 100 ms): The wait duration before initiating a retry attempt.
The rule delivery.timeout.ms >= linger.ms + request.timeout.ms must always hold. If a broker fails, the producer waits for the metadata refresh, discovers the new leader, and resends the batch, repeating until the delivery timeout expires.
Preventing duplicates and reordering
When a network packet drops after the broker commits a batch to disk, the producer cannot tell whether the write succeeded. If it simply resends the batch, two risks emerge:
- Duplication: The broker appends the same records twice.
- Reordering: If an earlier batch is retrying while a later batch succeeds, message order is inverted.
Enabling enable.idempotence=true fixes both problems. The producer assigns each batch an internal Producer ID and incrementing sequence numbers. The broker tracks sequence numbers per partition, transparently ignoring duplicate submissions and preserving order across retries.
Handling terminal failures
If the broker cluster remains unreachable and delivery.timeout.ms elapses, the producer halts retries and invokes the asynchronous callback:
producer.send(record, (metadata, exception) -> {
if (exception != null) {
logger.error("Failed to deliver record after timeout", exception);
}
});A short reminder on callback discipline:
- Always inspect the callback exception parameter to catch exhausted retries.
- Unhandled callback failures lead to silent message drops in upstream ingestion pipelines.