Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Data versioning

Pipelines & scenarios · Core Concepts Not Yet Covered

Data versioning

Mediumpipelines-67
data-versioningtime-travelsnapshotsreproducibilitylakefs

Question

What does data versioning mean, and how can you do it?

Solution

Data versioning means you can recover or reproduce what a dataset looked like at an earlier point in time, in the same way that Git lets you get an old version of code. It matters for audits, debugging, and reproducible machine learning.

Why you need it

  • A report from last quarter must be reproduced exactly, for an auditor.
  • A model was trained on data that has since changed. To retrain or explain it, you need that exact data.
  • A pipeline bug corrupted a table. You want to see it before the bug and roll back.
  • Someone asks why a number changed between Monday and Tuesday.

Table format time travel and snapshots

Delta Lake, Apache Iceberg and Hudi record each change as a new version of the table, with a log of which files belong to each version. You can query or restore an earlier version:

SELECT * FROM sales.orders VERSION AS OF 120;
SELECT * FROM sales.orders TIMESTAMP AS OF '2025-03-01 08:00:00';

Iceberg has snapshots, and tags or branches that give a name to a version (for example end_of_quarter_q1). Warehouses have their own versions of the idea, such as Snowflake Time Travel and BigQuery time travel and snapshots. These windows are limited by a retention setting, and old files are removed by cleanup jobs (VACUUM, snapshot expiry), so for permanent records you must keep a named snapshot or a copy.

Branches and Git-like tools

Tools such as lakeFS and Project Nessie add branching, commit and merge over data: you create a branch of the lake, run a pipeline against it, test the result, and merge only if it is good. That gives safe experiments and atomic multi-table publishes.

Versioned paths

A simple approach is to write each run's output to a versioned path (/curated/orders/run_date=2025-03-01/version=3/) and keep a pointer to the current one. It needs no special tool, and works for ML datasets and exports.

Practical points

  • Record the data version used by each model training run or report, alongside the code version.
  • Versions cost storage, so set retention deliberately.
  • Versioning the data is not a replacement for backups in a separate location.

Summarise it as: code has Git, data needs the equivalent, and open table formats give you much of it for free.

PreviousNext