Accepted answer
We saw the same issue, fixing the partition filter dropped runtime 60%.
Daily job reads ~50k small JSON event files, a few KB each, from S3. Listing plus reading takes 20+ minutes before any actual processing starts.
df = spark.read.json("s3a://bucket/events/2026/09/02/*")Compacting upstream isn't an option right now. What can we do on the read side?
Accepted answer
We saw the same issue, fixing the partition filter dropped runtime 60%.
Idempotent writes with merge keys saved us during backfills.
Increase fs.s3a.connection.maximum and parallelize the listing itself. Also check whether you're on the S3A committer v1 vs v2/magic, that affects listing overhead too.
A lightweight Lambda that compacts new files into larger chunks every hour, decoupled from the daily Spark job, fixes this at the source without touching upstream producers at all.
Sometimes the real fix is a product change so you stop needing that join at all.
Consider DuckDB or Polars for this size before spinning up a cluster.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Another path: push the compute to the warehouse if the data's already there.
Start with the execution plan, numbers beat guesses.
S3 listing is the bottleneck here, not the read itself. Use a manifest file with a pre-generated list of paths instead of a glob, so Spark skips the expensive S3 LIST calls.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.