Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Design a topic for order events

Kafka · Operations & Scenarios

Design a topic for order events

Hardkafka-61
system-designtopic-architecturee-commercescenario

Question

Design the Kafka topic(s) for an e-commerce order service. What decisions do you make?

Solution

Designing the Kafka topic architecture for an e-commerce order service requires balancing per-order message ordering, zero-data-loss durability, schema governance, and consumer group isolation. The foundation is a single partitioned topic named orders with messages keyed by order_id to guarantee that lifecycle state changes for any given order are consumed in chronological sequence. Durability is enforced using replication factor 3 with min.insync.replicas=2 and producer acks=all, supported by schema evolution via Schema Registry, a dead letter topic for malformed payloads, and separate consumer groups for fulfillment and analytics.

Topic structure and message keying

Use a single topic named orders containing an event_type header or payload attribute (OrderCreated, PaymentAuthorized, OrderShipped, OrderDelivered). Using a single topic prevents race conditions where downstream fulfillment workers read a PaymentAuthorized event from a separate payments topic before the OrderCreated event has arrived from an orders topic. Keying every record by order_id guarantees that all lifecycle updates for a specific customer order hash to the same partition, preserving causal ordering.

Sizing partitions for throughput

Estimate partition count based on peak holiday order volume:

  • If normal traffic is 1,000 orders per second, size for peak surges of 5,000 events per second.
  • If a single consumer instance can safely process 400 events per second against downstream database sinks, you need at least 13 partitions (5000 / 400).
  • Provisioning 16 partitions accommodates peak surges and leaves headroom for horizontal consumer scaling without repartitioning.

Replication, durability, and retention

Configure broker and producer settings for maximum durability:

  • Replication factor 3 across separate availability zones.
  • min.insync.replicas=2 combined with producer acks=all and enable.idempotence=true to guarantee zero message loss during broker failures.
  • Set retention to 14 days, providing ample time for downstream warehouses to recover from incidents and replay historical batches.
  • Enforce schema contracts using Avro or Protobuf with backward compatibility rules registered in Schema Registry.

Consumer isolation and dead letter handling

Separate downstream domains using independent consumer groups:

  • fulfillment-service-group: Real-time inventory and shipping workers.
  • analytics-lakehouse-group: Batched consumers streaming data into Iceberg or Snowflake.

Configure a dead letter topic named orders_dlq. If a consumer cannot deserialize or validate a corrupt order event, it writes the record to orders_dlq with failure headers and commits the partition offset, preventing pipeline blockages.

PreviousNext