Catch specific errors you know how to handle, let everything else fail loudly, and always clean up. In a pipeline, the worst outcome is a script that hits an error, hides it, and reports success.
Be specific
try:
amount = Decimal(row["amount"])
except (InvalidOperation, KeyError) as e:
...Catch the exact exceptions you expect. A bare except: or except Exception: pass also catches typos, KeyboardInterrupt and out-of-memory errors, and the job continues in a broken state.
Fail loudly so the orchestrator can act
If the job cannot do its work, let the exception propagate, or exit with a non-zero code. Airflow, Prefect and cron monitors all work from that signal: they retry, alert, and block downstream tasks. A script that prints "error" and exits with 0 looks like a success.
Decide per error: bad record or bad batch
- A single malformed row: send it to a quarantine file or table with the error reason, and carry on. Count them.
- Too many bad rows, a missing file, a failed connection, a schema change: fail the whole run. Set a threshold (for example more than 0.1 percent bad rows) so a flood of errors is treated as a batch failure.
Add context when you re-raise
try:
load_partition(path)
except Exception as e:
raise RuntimeError(f"load failed for {path}") from eraise ... from e keeps the original traceback and adds what you were doing, which saves a lot of time during an incident.
Clean up reliably
Use with blocks or finally for connections, files and temp directories, so they are closed even when something fails.
with get_connection() as conn:
...Logging
Log the identifiers (file name, batch id, row number), not the whole record, and never personal data or secrets. The log line should let someone find the failing input in a minute.
Retries
Retry only transient errors, a limited number of times. Permanent errors, such as a bad credential, should fail on the first attempt.
A good summary: be narrow in what you catch, loud in what you cannot handle, and explicit about what to do with bad data.