PySpark data engineering interview problem. Difficulty: beginner. Pattern: Broadcast Join. About 12 minutes. Part of the Pro drill bank.
Join orders to a small users table with F.broadcast. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Treat users as the small side. Join with F.broadcast(users) on user_id (inner). Select order_id, user_id, amount, name, email. Order by order_id. Assign result. The simulator prints a broadcast note; keep the habit for real clusters.
Input: orders + small users Output: order_id | user_id | amount | name | email 101 | 1 | 50.0 | Ada | a@x.com 102 | 2 | 75.0 | Alan | b@x.com 103 | 3 | 20.0 | Grace | c@x.com 105 | 1 | 60.0 | Ada | a@x.com Broadcast avoids shuffling the large side in production; result matches an inner join.
Topics: lakebench, pyspark, broadcast.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Join orders to a small users table with F.broadcast.
Treat `users` as the small side. Join with `F.broadcast(users)` on `user_id` (inner). Select `order_id`, `user_id`, `amount`, `name`, `email`. Order by `order_id`. Assign `result`. The simulator prints a broadcast note; keep the habit for real clusters.