Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Third-party API rate-limits us occasionally during a nightly pull. Currently retrying with a fixed 1-second sleep, which sometimes still triggers the limit repeatedly.
for attempt in range(3):
try:
resp = requests.get(url)
break
except requests.HTTPError:
time.sleep(1)What's a reasonable retry strategy here without pulling in a big dependency?
Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Document the grain decision, most BI bugs turn out to be grain bugs.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Prefer a staging table plus validation gate before promoting to prod tables.
Start with the execution plan, numbers beat guesses.
Note that merge on Delta still needs unique keys defined correctly.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Idempotent writes with merge keys saved us during backfills.
This matches our runbook for skewed keys.
Exponential backoff with jitter: sleep = base 2*attempt + random.uniform(0, base). Respect a Retry-After header if the API sends one.
How do you handle backfill without duplicating rows?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.