Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Join with Multiple Conditions

PySpark data engineering interview problem. Difficulty: beginner. Pattern: Joins. About 12 minutes. Part of the Pro drill bank.

Join on user_id then keep orders above the user's average. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Join orders to users on user_id, then keep rows where orders.amount > users.avg_order_amount. Select order_id, user_id, amount, name, avg_order_amount. Order by order_id. Assign result.

Constraints

  • Filter after the equi-join.
  • Order by order_id.

Examples

Input: orders vs avg_order_amount Output: order_id | user_id | amount | name | avg_order_amount 102 | 1 | 120.0 | Ada | 60.0 104 | 2 | 90.0 | Alan | 50.0 Equality join plus a residual inequality filter.

Topics: lakebench, pyspark, join, filter.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

Join with Multiple Conditions

Interview-style drill: Join on user_id then keep orders above the user's average.

Join `orders` to `users` on `user_id`, then keep rows where `orders.amount > users.avg_order_amount`. Select `order_id`, `user_id`, `amount`, `name`, `avg_order_amount`. Order by `order_id`. Assign `result`.