PySpark data engineering interview problem. Difficulty: advanced. Pattern: Broadcast Join. About 20 minutes. Part of the Pro drill bank.
Use F.broadcast join instead of a broadcast dict + UDF lookup. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Join orders to F.broadcast(users) on user_id. Select order_id, user_id, amount, name. Order by order_id. Assign result. Avoid broadcast variables + Python UDFs.
Input: orders + users Output: order_id | user_id | amount | name 101 | 1 | 50.0 | Ada 102 | 2 | 75.0 | Alan 103 | 1 | 20.0 | Ada 104 | 3 | 40.0 | Grace Broadcast join replaces dict UDF lookups.
Topics: lakebench, pyspark, broadcast.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Use F.broadcast join instead of a broadcast dict + UDF lookup.
Join `orders` to `F.broadcast(users)` on `user_id`. Select `order_id`, `user_id`, `amount`, `name`. Order by `order_id`. Assign `result`. Avoid broadcast variables + Python UDFs.