In our case the root cause was an implicit cast preventing pushdown.
Trying to parallelize a per-file transform across 8 workers with Pool.map over a list of pandas DataFrames already loaded in memory. The job hangs indefinitely after a few chunks.
with Pool(8) as pool:
results = pool.map(transform, dataframes)No error, no CPU usage, it's just stuck. What am I missing?
In our case the root cause was an implicit cast preventing pushdown.
In our case the root cause was an implicit cast preventing pushdown.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
This matches our runbook for skewed keys.
Prefer a staging table plus validation gate before promoting to prod tables.
Sometimes the real fix is a product change so you stop needing that join at all.
Pickling large DataFrames to ship to worker processes is expensive and can even deadlock on some platforms if a worker dies mid-transfer. Check memory pressure first, this often looks like a hang but is actually swapping.
If this is CPU-bound and the data fits, Polars' native multithreading might get you the speedup without manual multiprocessing at all.
Agree on the staging table swap. Atomic promote prevented partial reads.
ELI5 version: it's not broken, it's just slow because it's checking way more stuff than it needs to. Narrowing what it checks is almost always the fix.
Sometimes the real fix is a product change so you stop needing that join at all.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.