Sizing partitions begins by dividing your target peak throughput by the slower of your single-partition producer or consumer processing throughput. Because each partition can be assigned to only one consumer within a consumer group, the partition count sets the hard ceiling on horizontal consumer parallelism. You should size partitions with future traffic growth in mind because expanding partition counts later alters key-to-partition hashing, and Kafka never allows decreasing partition counts.
Sizing formula and consumer limits
Calculate the partition count using the slower processing link:
Partitions = Max( Target Peak Ingestion / Single-Partition Producer Rate, Target Peak Ingestion / Single-Partition Consumer Rate )
A concrete scenario helps illustrate this calculation:
- If a clickstream topic must ingest 80 MB/s, and a single producer can push 40 MB/s per partition, producer requirements suggest 2 partitions.
- If downstream consumer database writes can only process 10 MB/s per core, consumer requirements dictate 8 partitions (80 / 10).
- Setting the topic to 12 or 16 partitions leaves head room for traffic spikes and provides room to scale out consumer pods.
The irreversible partition constraint
You can increase partitions on an existing topic using administrative commands, but you can never reduce partitions on a topic. The default producer partitioner maps keys using a hash modulo operation against partition count. If you change the partition count from 8 to 16, subsequent writes for an existing customer ID will hash to a completely different partition, splitting that entity's historical event order across two partitions.
The overhead of excessive partitions
Creating thousands of partitions per broker creates real infrastructure costs:
- Each partition open segment requires open OS file descriptors and socket buffers.
- Replicating partitions consumes broker memory and internal metadata replication bandwidth.
- During broker restarts, recovering thousands of partitions extends leader failover times.
Aim for a sensible partition count that supports peak consumer concurrency over the next 12 to 18 months rather than over-provisioning hundreds of unused partitions.