The requirement is that the same feature is available in two forms: a big historical table for training, and a fast lookup for live predictions. The danger is that the two get computed differently, which causes training-serving skew.
raw events -> feature transformations (one codebase)
|-> offline store (warehouse/lake) -> training sets
|-> online store (Redis/Bigtable) -> model servingOffline store
Features are kept in the warehouse or lake, with an entity key (customer_id) and a timestamp for when each value was true. Training needs point-in-time correct joins. If a label is "customer churned on March 1", the features must be the values as of the day before, and never anything computed afterwards. Joining the latest feature values to old labels leaks the future into training, and the model looks great offline and fails in production. A feature store does this as-of join for you, and without one you write it carefully.
-- latest feature value at or before each label time
SELECT l.customer_id, l.label_ts, f.orders_30d
FROM labels l
LEFT JOIN features f
ON f.customer_id = l.customer_id AND f.feature_ts <= l.label_ts
QUALIFY ROW_NUMBER() OVER (PARTITION BY l.customer_id, l.label_ts
ORDER BY f.feature_ts DESC) = 1;Online store
A key-value store (Redis, Bigtable, DynamoDB) holding the latest value per entity, read in a few milliseconds when the model is called. A job pushes fresh values in, either streaming for features like "clicks in the last 10 minutes" or on a batch schedule for slow-moving ones.
One definition of each feature
The biggest source of skew is writing the logic twice, SQL for training and Java or Python for serving. Define each feature once, and have the platform (Feast, Vertex AI Feature Store, Databricks Feature Store, Tecton) materialise it to both stores. Test that the values match for a sample of entities.
Operations
- Backfill: when you add a feature, compute its history so older training examples can use it.
- Freshness monitoring: alert when the online values are older than the allowed lag.
- Drift: compare live feature distributions with training distributions.
- Ownership, documentation and versioning for each feature, so teams can reuse them.
Trade-off
A full feature store is heavy for one model. For a single team, start with a well-tested offline table plus a small online cache, and adopt a feature store when several models share features.