The interview room
Design an event-driven order processing system for food delivery. Kafka carries state transitions like OrderPlaced, OrderAccepted, DriverAssigned. You must use the Outbox pattern so DB and Kafka stay consistent, and timeout events so unaccepted orders cancel cleanly.
Interviewers love this prompt because it mixes a state machine, messaging reliability, and real failure modes (restaurant never accepts, driver cancels, dual-write bugs).
First, the basics
What "event-driven" means here
Services react to facts ("order placed") instead of calling a giant synchronous orchestrator for every hop. The order service writes durable state. Other services (restaurant, dispatch, notifications) consume events and act.
The dual-write problem
Naive flow:
1. INSERT order into Postgres. 2. produce OrderPlaced to Kafka.
If you crash between 1 and 2, the DB has an order and nobody else knows. If you reverse the order and crash after Kafka, you have a phantom event with no DB row. Two non-atomic writes are the bug. This is the dual-write problem.
Outbox pattern (the fix)
In one database transaction:
- Write the order row (or state transition).
- Write an outbox row with the event payload.
A separate publisher reads outbox rows and produces to Kafka, then marks them published. DB commit is the atomic boundary. Kafka delivery becomes reliable asynchronous propagation.
State machines
Orders are not free-form strings. Legal transitions matter:
Placed --> Accepted --> DriverAssigned --> PickedUp --> Delivered
\ \ /
+-----> Cancelled <----+Illegal jumps (Delivered → Placed) must be rejected. Timeouts are first-class events (OrderAcceptanceTimedOut) that drive cancels.
Idempotent consumers
At-least-once delivery means consumers may see duplicates. Processing OrderPlaced twice must not charge twice or assign two drivers. Key on event_id or (order_id, transition).
The problem you are designing
Food delivery order flow:
- Persist order state transitions reliably.
- Publish events without critical loss from dual-write races.
- Timeout unaccepted orders.
- Downstream consumers update logistics and notifications.
- Regional scale with meal-time bursts.
What good looks like
You name dual-write as the core reliability bug. You draw outbox. You sketch a state machine. You add a timeout path. You mention idempotent consumers. You discuss conflict cases (driver assigned, then restaurant cancels).
Clarifying questions
- Who is SoR for order state: the order DB, or a pure event-sourced log?
- Timeout durations for restaurant accept and driver assign?
- Sync response to the customer app vs eventually consistent confirmation?
- Multi-city topics vs one big cluster?
- Payment authorization timing relative to
OrderPlaced?
Constraints
- Meal peaks create bursty produce/consume.
- Losing
OrderPlacedis a Sev-1 for restaurants and customers. - Duplicate notifications are annoying; duplicate charges are catastrophic.
Out of scope
- Full GIS routing algorithm for drivers.
- Detailed payment network design (mention authorize/capture at high level).
- Mobile app UX beyond API contracts.
Why food delivery is a great prompt
Orders cross organizational boundaries. The restaurant is not your database. The driver is not your database. Notifications are best-effort. Payments are strict. That mix forces you to talk about reliability classes, not one generic "event bus."
Dual-write in slow motion
Walk this with the interviewer:
1. API handler begins. 2. INSERT into orders succeeds. 3. Process is OOM-killed before producer.send. 4. Customer sees "order placed" from the API response that read the DB. 5. Restaurant never receives OrderPlaced. 6. Support invents a manual cancel thirty minutes later.
That story is worth more than naming five Kafka features.
Outbox mental model
Treat the outbox table as "events the business has decided are true but the world has not heard yet." The publisher is a reliable postman, not the decision maker. Decision making stays in the transactional order service.
Timeout as a first-class citizen
Without timeouts, Placed becomes a junk drawer. Restaurants forget. Customers wait. Inventory holds linger. A timeout event is not a niche feature. It is how you keep the state machine honest when humans do not respond.
What good looks like beyond the diagram
Mention idempotent consumers, per-order partitioning keys, outbox lag alerts, and one conflict policy (cancel after assign). If you only draw boxes, they will keep asking "what if" until those appear.
Reliability classes (teach explicitly)
Not every side effect needs the same guarantee:
- Order state + payments: hard correctness, idempotent, audited
- Restaurant ticket: durable, delayed is bad but recoverable
- Push notification: best effort with debounce
- Analytics copy of events: at-least-once into the lake is fine
Naming reliability classes prevents over-engineering notifications and under-engineering payments.
Burst and backpressure awareness
At dinner peak, dispatch consumers may lag. Kafka retains. You should say what the lag SLO is and when you page. Backpressure belongs in this design even if the prompt emphasizes outbox.