Accepted answer
Also add a DLQ on the Lambda's trigger so a failed crawler-start call doesn't just silently vanish. You want visibility when this specific step fails.
A Lambda kicks off a Glue crawler after new files land, but occasionally the crawler start call itself times out (Lambda's 15-minute max), leaving things in an unclear state.
Is Lambda even the right tool to babysit a crawler run?
Accepted answer
Also add a DLQ on the Lambda's trigger so a failed crawler-start call doesn't just silently vanish. You want visibility when this specific step fails.
Careful with NULL in join keys, they'll drop rows in an inner join.
We saw the same issue, fixing the partition filter dropped runtime 60%.
We were polling GetCrawler in a loop inside the Lambda, which is exactly the anti-pattern. Switching to EventBridge.
Idempotent writes with merge keys saved us during backfills.
This matches our runbook for skewed keys.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In our case the root cause was an implicit cast preventing pushdown.
Lambda should only start the crawler async and return immediately, don't have it wait or poll for completion. Use EventBridge on the Glue crawler's state-change event to trigger whatever needs to run after.
Start with the execution plan, numbers beat guesses.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.