Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Feature Store for ML (Batch + Real-Time)

HardPro55 min read

Design a feature store for batch training (millions of vectors) and real-time inference (<10ms), with point-in-time correctness.

system-designfeature-storemlstreamingpoint-in-time

Interview framing

You are designing a feature store for machine learning in a senior-leaning DE interview. About 55 minutes. The scoring center is training-serving skew and point-in-time correctness, not only Redis plus Snowflake.

The prompt:

Design a feature store that serves batch training (millions of historical feature vectors) and real-time inference (fetch one feature vector in under 10 ms), with point-in-time correctness so training does not leak future information.

Background from first principles

A feature is an input signal to a model: average spend last 30 days, clicks in the last hour, account age. Training needs large historical tables. Inference needs a single entity lookup in milliseconds.

If training computes average spend last 30 days from a warehouse dump that already includes today's purchase, the model learns the future. Offline metrics look great. Online performance collapses. That failure mode is training-serving skew (and label leakage when the label itself sneaks into features).

A feature store exists so one definition of a feature materializes into two serving surfaces: offline (lakehouse) and online (low-latency key-value), without two conflicting SQL dialects living in different teams' notebooks.

Contrast with analytics dashboards. Dashboards can tolerate approximate tiles and human judgment. Models amplify silent mismatches. That is why this problem is stricter about consistency.

Expanded problem and constraints

  • Offline: read millions to billions of rows for training dataset generation.
  • Online: get(entity_keys) under about 10 ms p99 in the inference region.
  • Point-in-time correct joins for labels at time T.
  • Backfill after feature logic changes.
  • Monitor training-serving skew (compare offline recompute vs online values).
  • Replayable compute from raw events or tables.

What good looks like

You tell a skew story first. You propose two stores, one definition. You explain point-in-time joins with a tiny timeline example. You make the online path a KV lookup, not a warehouse query. You discuss missing-feature behavior and versioning.

You also separate batch-daily features from true streaming features instead of pretending everything is real-time on day one.

Clarifying questions

  • Entity keys: user, device, listing, or multiple entity types?
  • Do online features need time travel or only latest values?
  • Batch-only features (T+1) vs true streaming features (seconds)?
  • Who authors feature definitions: ML or DE?
  • Inference QPS and payload size per vector?
  • Multi-region inference?
  • What happens when a feature is missing at score time?

Scale and estimation prompts

  • Offline storage is cheap at lakehouse scale; online keeps latest features per entity (plus short history only if required).
  • Inference QPS may be thousands; latency dominates. Redis/Dynamo/Cassandra-style stores.
  • Feature count per model: dozens to hundreds. Watch online memory and network.
  • Estimate vector size: 200 features x 8 bytes is small; nested maps and huge embeddings change the story.

Out of scope for v1

  • Using the warehouse as the online inference path.
  • Guaranteeing identical streaming freshness for every feature on day one (mix batch and stream deliberately).
  • Full time travel for every online key unless product requires it.
  • Building a complete managed platform clone if a thin in-house spine meets the latency and PIT needs.
  • Training the model itself (focus on feature serving and correctness).

Why DE owns part of this

Feature stores sit between data platform and ML platform. Data engineers bring correctness under time, backfills, and reliable materialization. ML engineers bring feature intent and model constraints. The interview rewards candidates who can speak both languages without dissolving ownership.

Extra product constraints

Confirm whether fraud models need sub-second streaming features while growth models can wait for daily batch. Confirm embedding features and their sizes. Confirm whether point-in-time correctness is required only for training or also for batch scoring jobs that simulate past decisions.

What interviewers listen for

Skew story, PIT join, two stores one definition, sub-10ms KV path, versioning, and monitoring. Logo soup without those themes fails.