Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Observability for pipelines: what to measure

Pipelines & scenarios · Core Concepts Not Yet Covered

Observability for pipelines: what to measure

Mediumpipelines-61
observabilitymonitoringmetricslineagelogging

Question

What metrics and logs do you need to operate pipelines well?

Solution

To run pipelines well you need to see three things: did it run, is the data right, and what did it cost. That means collecting run metrics, data metrics and cost metrics, with logs you can search.

Run health

  • Status of each run, and its duration compared to its usual duration (a trend line catches slow growth).
  • Retries and failure counts, queue wait time, and SLA misses.
  • Resource use: CPU, memory, spill, and shuffle sizes for Spark; slot time or credits for warehouses.

Data health

  • Rows in and rows out at each step, because a step that suddenly outputs 40 percent fewer rows is an incident that no error message announces.
  • Freshness: the age of the newest data in each output table, compared to its promise.
  • Data quality test results: pass or fail counts, with a history.
  • Schema changes detected, bad record rates, and quarantine volumes.

Cost

Credits, bytes scanned or DBUs per pipeline, per team, and per run, tagged so you can attribute them. Cost spikes are often the first sign of a logic change.

Logs

Use structured logs (JSON) with a run id, task name, table name and batch date on every line, so you can find all the lines for one run across systems. Keep them where people can search (not only on a cluster that gets deleted). Log counts and decisions, not personal data.

Lineage

Know which tables feed which, so when something fails you can list everything downstream (the blast radius) and tell the owners. Catalogs and tools like OpenLineage capture this automatically.

Dashboards and alerts

Build one overview of all pipelines: green or red, last success, freshness, and trend. Alerts should have an owner, a severity, and a link to the runbook and the logs. Put the most important metrics (freshness and SLA) first.

Priority for a small team

Start with run status, duration, freshness and row counts. Add quality tests and cost next, and lineage as you grow. Say that you would measure what you will act on, and avoid collecting numbers that nobody looks at.

PreviousNext