Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Reprocess after a bug in transformation logic

Pipelines & scenarios · Production Scenarios

Reprocess after a bug in transformation logic

Hardpipelines-53
scenariobugbackfillincidentpostmortem

Question

You discover a transformation bug that has been producing wrong values for 3 weeks. Walk through your response.

Solution

Treat it as an incident with a clear order: understand the impact, tell the people affected, fix the logic with a test, repair the data safely, check it, and learn from it. Speed matters less than accuracy, because a second wrong fix makes things worse.

1. Assess the impact

Define exactly what was wrong (which column or metric, which condition), when it started (find the first bad partition, using code history and data comparison), and how large the error is (a 0.5 percent rounding issue and a doubled revenue are different incidents). List the consumers using lineage: dashboards, reports sent to clients or finance, exports, ML features and models trained on this data, and downstream tables.

2. Communicate early

Tell the owners of those consumers now, even before the fix is ready: what is wrong, since when, what the numbers should look like roughly, and when you will update them. For anything used in decisions or sent outside the company, involve the business owner at once. People trust a team that reports its own mistake quickly.

3. Fix the code with a test

Write a test that reproduces the bug with a small input, watch it fail, then fix it and watch it pass. That test prevents the bug from returning. Review the fix with a colleague.

4. Repair the data

Rebuild the affected partitions with the corrected logic, using idempotent overwrites, in a staging area or with a clone if possible, and check before publishing. Keep the old values in a backup table for comparison. Process in date order, and rebuild downstream models after the source tables are right.

5. Validate

Compare old and corrected values by day, and make sure the differences are what the bug explains, and nothing else changed. Reconcile with an independent source (the finance ledger, the source system).

6. Close the loop

  • Tell consumers the corrected numbers are live, with the size of the change by period.
  • For models trained on the bad data, tell the ML owners so they can decide whether to retrain.
  • Write a blameless postmortem: what happened, why tests did not catch it, how long it ran, and the follow-up actions (a new data test, a reconciliation check, a review rule).

Add a detection step

Three weeks is a long time. Ask what monitoring would have shown it sooner, such as a comparison of totals with the source or a drift alert on key metrics, and add it.

PreviousNext