Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Lakeflow Declarative Pipelines (DLT)

Snowflake, BigQuery & Databricks · Databricks

Lakeflow Declarative Pipelines (DLT)

Mediumwarehouses-43
databricksdeclarative-pipelinesdltlakeflowcdc

Question

What are Delta Live Tables / Lakeflow Declarative Pipelines?

Solution

Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables, or DLT) is a framework where you declare the tables you want, as queries, and the platform runs them. It figures out the order, manages the checkpoints, retries and infrastructure, and tracks data quality.

What you write

CREATE OR REFRESH STREAMING TABLE bronze_orders
AS SELECT * FROM STREAM read_files('s3://landing/orders/', format => 'json');

CREATE OR REFRESH MATERIALIZED VIEW daily_revenue
AS SELECT order_date, SUM(amount) AS revenue
   FROM silver_orders GROUP BY order_date;

A streaming table processes new data incrementally. A materialized view is kept up to date with the results of a query, incrementally where possible. The same can be written in Python with decorators.

What the framework does for you

  • Works out dependencies from the queries, and runs them in the right order.
  • Manages streaming checkpoints, retries failed steps, and scales the compute.
  • Lets you run in triggered mode (process and stop) or continuous mode.
  • Records an event log with data quality results and lineage.
  • Applies data quality rules called expectations (see the next question).

Change data capture

The AUTO CDC API (previously called APPLY CHANGES) takes a stream of changes and applies them to a target table as SCD Type 1 (overwrite) or SCD Type 2 (keep history), handling out-of-order events through a sequence column. Writing this by hand with MERGE is long and easy to get wrong. The older APPLY CHANGES name still works.

Naming

Databricks renamed Delta Live Tables to Lakeflow Declarative Pipelines in 2025, and the core declarative pipeline engine was contributed to Apache Spark as Spark Declarative Pipelines. Because the names changed, you will see all of them in job postings and docs, so say that you know they refer to the same idea.

Trade-offs

You give up some control for simplicity, and you must follow the framework's rules, such as not mixing arbitrary side effects into pipeline code. For straightforward bronze, silver and gold layers, it removes a lot of boilerplate.

PreviousNext