Accepted answer
Push-down predicate only helps if the catalog partition filter is actually applied. Log the input file list count before and after the predicate to check.
Migrated a PySpark script from the DataFrame API to Glue DynamicFrame for catalog integration. Same logic, but DPU hours went from around 12 to 28 per run.
Using bookmark and push_down_predicate on the partition column. Anything obvious that makes DynamicFrame more expensive?
Accepted answer
Push-down predicate only helps if the catalog partition filter is actually applied. Log the input file list count before and after the predicate to check.
We saw the same issue, fixing the partition filter dropped runtime 60%.
We stayed on DataFrames for transforms and used Glue only for catalog read/write boundaries. Cut cost about 40% doing that.
In our case the root cause was an implicit cast preventing pushdown.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
DynamicFrame conversions and ResolveChoice can force extra passes. Profile with the Spark UI, there's often an unexpected .toDF()/.fromDF() round trip hiding per branch.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
We replaced custom sensors with data contracts and row count checks.
Document the grain decision, most BI bugs turn out to be grain bugs.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Worth measuring the serialized size before choosing broadcast.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.