@charan_reddy
Pro member
Consultant helping teams ship reliable Iceberg workloads.
Trying to parallelize a per-file transform across 8 workers with Pool.map over a list of pandas DataFrames already loaded in memory. The job hangs indefinitely after a few chunks. No error, no CPU usage, it's just stuck. What am I missing?
This worked in dev on sample data but fails at full volume. Details: Context: Covering index worth it for wide dimension table? Happy to share schema snippets or metrics if useful.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.