Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Manual vs auto offset commit

Kafka · Operations & Scenarios

Manual vs auto offset commit

Mediumkafka-52
offset-commitauto-commitmanual-commitat-least-once

Question

Should you use enable.auto.commit? When do you commit manually?

Solution

Automated offset commits should rarely be used in mission-critical pipelines because enable.auto.commit commits offsets on an elapsed timer regardless of whether your code has actually finished processing the records. This behavior risks silent data loss if an application crashes midway through processing a batch, or causes duplicate processing if a failure occurs before the timer fires. Production systems use manual offset commits after verifying downstream processing, using synchronous commits for strict durability or asynchronous commits for high throughput combined with a final synchronous commit during graceful shutdown.

Why auto-commit causes data loss

When enable.auto.commit=true (with auto.commit.interval.ms=5000), offset commits happen automatically inside subsequent poll() invocations based on wall-clock time:

  • A consumer polls 500 records and begins executing downstream transformations.
  • Five seconds pass while the application is processing record 200.
  • If the application crashes at record 250, the offsets up to record 500 may already have been committed during an internal timer tick.
  • When the container restarts, it skips records 251 through 500 entirely, causing permanent, unrecoverable data loss.

Choosing between commitSync and commitAsync

Manual offset control puts the commit point squarely under developer control:

  • commitSync: Blocks the consumer thread until the broker confirms the offset write. It offers simple error handling but introduces network round-trip latency after every batch.
  • commitAsync: Dispatches offset commits asynchronously without blocking the main loop, maximizing ingestion throughput. If an asynchronous commit fails due to a transient network glitch, a subsequent successful commit of a higher offset naturally supersedes it.

A standard production pattern uses commitAsync inside the main message processing loop, wrapped in a try-finally block that calls commitSync during application shutdown to commit the final offsets cleanly.

Commit granularity rules

Always commit offsets per batch, never per record. Committing an offset after every single record swamps the __consumer_offsets topic with network requests, causing broker thread exhaustion and cratering consumer throughput.

PreviousNext