Treat it as an incident with a clear order: understand the impact, tell the people affected, fix the logic with a test, repair the data safely, check it, and learn from it. Speed matters less than accuracy, because a second wrong fix makes things worse.
1. Assess the impact
Define exactly what was wrong (which column or metric, which condition), when it started (find the first bad partition, using code history and data comparison), and how large the error is (a 0.5 percent rounding issue and a doubled revenue are different incidents). List the consumers using lineage: dashboards, reports sent to clients or finance, exports, ML features and models trained on this data, and downstream tables.
2. Communicate early
Tell the owners of those consumers now, even before the fix is ready: what is wrong, since when, what the numbers should look like roughly, and when you will update them. For anything used in decisions or sent outside the company, involve the business owner at once. People trust a team that reports its own mistake quickly.
3. Fix the code with a test
Write a test that reproduces the bug with a small input, watch it fail, then fix it and watch it pass. That test prevents the bug from returning. Review the fix with a colleague.
4. Repair the data
Rebuild the affected partitions with the corrected logic, using idempotent overwrites, in a staging area or with a clone if possible, and check before publishing. Keep the old values in a backup table for comparison. Process in date order, and rebuild downstream models after the source tables are right.
5. Validate
Compare old and corrected values by day, and make sure the differences are what the bug explains, and nothing else changed. Reconcile with an independent source (the finance ledger, the source system).
6. Close the loop
- Tell consumers the corrected numbers are live, with the size of the change by period.
- For models trained on the bad data, tell the ML owners so they can decide whether to retrain.
- Write a blameless postmortem: what happened, why tests did not catch it, how long it ran, and the follow-up actions (a new data test, a reconciliation check, a review rule).
Add a detection step
Three weeks is a long time. Ask what monitoring would have shown it sooner, such as a comparison of totals with the source or a drift alert on key metrics, and add it.