Automated offset commits should rarely be used in mission-critical pipelines because enable.auto.commit commits offsets on an elapsed timer regardless of whether your code has actually finished processing the records. This behavior risks silent data loss if an application crashes midway through processing a batch, or causes duplicate processing if a failure occurs before the timer fires. Production systems use manual offset commits after verifying downstream processing, using synchronous commits for strict durability or asynchronous commits for high throughput combined with a final synchronous commit during graceful shutdown.
Why auto-commit causes data loss
When enable.auto.commit=true (with auto.commit.interval.ms=5000), offset commits happen automatically inside subsequent poll() invocations based on wall-clock time:
- A consumer polls 500 records and begins executing downstream transformations.
- Five seconds pass while the application is processing record 200.
- If the application crashes at record 250, the offsets up to record 500 may already have been committed during an internal timer tick.
- When the container restarts, it skips records 251 through 500 entirely, causing permanent, unrecoverable data loss.
Choosing between commitSync and commitAsync
Manual offset control puts the commit point squarely under developer control:
- commitSync: Blocks the consumer thread until the broker confirms the offset write. It offers simple error handling but introduces network round-trip latency after every batch.
- commitAsync: Dispatches offset commits asynchronously without blocking the main loop, maximizing ingestion throughput. If an asynchronous commit fails due to a transient network glitch, a subsequent successful commit of a higher offset naturally supersedes it.
A standard production pattern uses commitAsync inside the main message processing loop, wrapped in a try-finally block that calls commitSync during application shutdown to commit the final offsets cleanly.
Commit granularity rules
Always commit offsets per batch, never per record. Committing an offset after every single record swamps the __consumer_offsets topic with network requests, causing broker thread exhaustion and cratering consumer throughput.