Producer throughput and network efficiency are primarily governed by combining individual messages into compressed record batches in client memory. The configuration batch.size sets the maximum byte threshold for a partition batch, while linger.ms introduces a small delay to give incoming records time to pack together before transmission. Applying compression.type compresses these batches on the client, achieving significantly better compression ratios on larger batches while buffer.memory caps the producer total memory pool.
How batch size and linger work together
When a producer calls send(), records accumulate in memory buffers partitioned by topic and partition:
- batch.size (default 16 KB): The target memory size for an individual partition batch. For high-volume pipelines, increasing this to 64 KB or 128 KB allows larger payloads to travel together.
- linger.ms (default 0 ms): The maximum time the producer waits for additional records before dispatching an incomplete batch.
With linger.ms=0, the producer dispatches immediately as soon as a network thread is free. Setting linger.ms=10 instructs the producer to wait up to 10 milliseconds. If high traffic fills batch.size within 2 milliseconds, the batch sends immediately. If traffic is light, waiting 10 milliseconds allows multiple messages to coalesce, turning dozens of tiny network round-trips into a single efficient transfer.
Compression algorithms and batch size synergy
Compressing individual messages is inefficient because short strings lack repeating patterns. Kafka compresses the entire batch at once:
- lz4 and snappy: Offer balanced compression ratios with ultra-low CPU overhead, making them standard for real-time streaming.
- zstd: Delivers superior compression ratios for JSON or structured logs at the cost of slightly higher CPU usage.
- Larger batches provide richer text patterns, drastically improving compression ratios and saving network bandwidth.
Handling buffer memory exhaustion
All in-flight partition batches share a single memory pool controlled by buffer.memory (default 32 MB). If downstream brokers slow down and this memory pool fills up, the producer blocks subsequent send() calls up to max.block.ms before throwing a TimeoutException.