A single bad record should never be able to stop 50 million good ones, and it should never vanish quietly either. So parse tolerantly, set the bad rows aside with the reason, and decide with a threshold whether the batch as a whole is acceptable.
Parse permissively
Instead of failing at the first malformed row, read the data in a mode that keeps going. In Spark, the permissive mode keeps the raw text of rows that do not parse in a corrupt-record column. For custom code, wrap per-row parsing in a try block, and collect the failures.
Quarantine, with a reason
Write every bad record to a separate quarantine table, with the original raw text, the error reason (for example "amount not numeric"), the source file, the line or offset, and the run id. This lets someone investigate, and lets you replay it later.
quarantine_orders: raw_record | error_reason | source_file | line_no | run_id | ts
Set a threshold
Not every bad rate is acceptable. Decide a rule such as "fail the run if more than 0.1 percent of rows are bad", or "fail if any row is missing a primary key". A handful of bad rows out of 50 million is a data issue to follow up, while 20 percent bad rows means the source format changed, and loading the rest would produce nonsense. The threshold turns a quiet degradation into a loud failure when it matters.
Alert the owners
Send a summary to the source owner: how many rows were rejected, what types of errors, with examples. Many issues, such as an upstream bug that writes a date in two formats, are fixed in a day when the owner sees the evidence.
Replay
After the data or the parser is fixed, reprocess the quarantined rows, and merge them into the main table. Because the load is idempotent (keyed MERGE), replaying cannot duplicate.
What not to do
Do not silently drop bad rows, because you will lose data and never know. Do not keep failing the whole batch for one row either, because the pipeline becomes fragile and your team gets paged for something that needs a one-line fix elsewhere. Also track the bad rate over time, since a slow rise is an early warning.