Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Real-Time Fraud Detection (<100ms Latency)

HardFree75 min read

Design fraud detection for ~5k TPS payments with <100ms decision latency, including features, scoring, and feedback loops.

system-designstreamingfraudlow-latencyml

The interview setup

You are in a mid-level data engineering interview. The prompt sounds like machine learning platform work. The interviewer is actually testing something sharper: can you protect a synchronous payment path with a hard latency budget, while still building a learning loop that improves over time?

A customer taps a card at a coffee shop. Your payment API needs an approve, decline, or step-up answer in under 100 milliseconds at about 5,000 transactions per second. That is not "run Spark overnight and email Risk." That is a real-time decision service with an asynchronous training world wrapped around it.

What interviewers reward here:

  • You feel why 100ms changes the design
  • You separate the online decision path from the offline training path
  • You budget milliseconds on the whiteboard
  • You name fallbacks when features or models are unhealthy
  • You talk about delayed labels (chargebacks), not magically instant truth

Background concepts (learn these before drawing boxes)

What online fraud detection actually is

Fraud detection is a decision about a payment before money fully moves, or before capture completes. Inputs are the payment fields plus features: signals about the card, device, merchant, velocity, geography, and history.

Typical outputs:

  • Approve - let the payment through
  • Decline - hard stop
  • Review / step-up - OTP, 3-D Secure, or a human queue

You are not building a BI dashboard. You are building a gate on the money path. Product, finance, and risk all care, and they disagree about false positives.

Why latency is a product requirement, not a nice-to-have

Card networks and checkout UX have tight timeouts. If fraud takes 800ms, either payments fail or product removes your check. Healthy fraud platforms therefore split into layers:

1. Online path - tiny latency budget, nearby dependencies only 2. Nearline path - enrichment and heavier features that update profiles asynchronously 3. Offline path - training, evaluation, shadow scoring, rule analytics

If you remember one sentence for the rest of your career: never put a batch warehouse query on the authorization critical path.

Online vs offline feature stores

A feature is a computed signal used by rules or models: transaction count in the last hour, average ticket over 30 days, whether the device was seen before, merchant fraud rate over 7 days.

An online feature store serves precomputed or fast-computed features with low latency. Think Redis, Cassandra, DynamoDB, or a purpose-built feature service.

An offline feature store / training set stores historical feature values joined to labels for model training.

The interview killer is training-serving skew: you train on features that do not exist online, or that are computed differently online. The notebook looks brilliant. Production is blind.

Rules vs models

Mature stacks combine:

  • Rules - explicit, auditable, fast ("decline if country flip + high amount + new device")
  • ML models - score risk from many features
  • Policy - model score + rules + business thresholds become the final action

Rules are guardrails and kill switches. Models catch patterns rules miss. Strong candidates put rules as overrides, not as an afterthought.

Feedback loops and delayed labels

Truth for fraud is often late. A chargeback may arrive days or weeks after purchase. Your training pipeline must join decisions to outcomes by payment id with lag. Do not pretend labels arrive at decision time.

Human review outcomes and merchant reports are also labels, with different delay and noise. Good designs log the feature vector and model version at decision time so later labels can attach cleanly.

The expanded problem

Design fraud detection for a payment processor at roughly 5,000 TPS. Return a decision in under 100ms p99. Include feature fetch or computation, model scoring and rules, decision logging for audit and training, a feedback loop for retraining, safe rollouts, and kill switches.

Assume you own the decision service. Kafka (or similar) is fine for async logs. Object storage and a warehouse are fine for training. They are not fine on the sync path.

Customer / merchant checkout
        |
        v
   Payment API ----sync----> Fraud Decision Service ----> approve | decline | step-up
                                      |
                                      +--> online feature store
                                      +--> model + rules + policy
                                      |
                                      v
                               async decision log (Kafka)
                                      |
                                      v
                         lake + label join + retrain + shadow

Constraints and scale prompts (say numbers out loud)

  • 5,000 TPS sustained is about 432 million transactions per day; size for peak, not average
  • Latency budget is scarce; feature store p99 and network hops dominate
  • Model location matters: embedded in-process vs remote model server in the same AZ
  • Idempotency: clients retry; the same payment id must not silently flip decisions

Example latency budget you can write on the board:

  • Network into decision service: 5–10ms
  • Feature fetch (batched): 20–40ms
  • Model inference: 5–20ms
  • Rules and policy: 5–10ms
  • Margin for GC and jitter: the rest

If feature fetch alone is 80ms, the design is already in trouble.

What good looks like in the room

A strong answer: 1. Clarifies hard decline vs step-up vs async review 2. Draws sync path and async training path as two systems that meet at a log 3. Breaks the 100ms budget explicitly 4. Names online feature store plus a degradation policy 5. Explains delayed labels and shadow or canary model deploy 6. Mentions audit trails for declines and a kill switch for bad models

Clarifying questions to ask

  • Is 100ms p99 end-to-end from the payment API, or fraud service internal only?
  • Which outcomes are allowed: hard decline, step-up, human review?
  • What false positive rate is tolerable? Declining good customers kills conversion.
  • Do we already have an online feature store, or design from scratch?
  • Single region or multi-region active-active payments?
  • Primary label source: chargeback, manual review, or both?
  • Regulatory audit needs for decline reasons?

Out of scope (say this out loud)

  • Inventing the "best" neural net architecture from scratch
  • A full PCI deep dive (mention awareness, do not boil the ocean)
  • Building the entire card network
  • Real-time graph neural nets unless the interviewer invites the stretch

Your job is the decision path and the data feedback loop, not a research paper.