PySpark data engineering interview problem. Difficulty: beginner. Pattern: Filtering. About 12 minutes. Free to practice.
Keep purchase rows with amount > 100 and project user_id, amount. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From df, keep rows where amount > 100 and event_type == 'purchase'. Select only user_id and amount. Order by user_id, then amount. Assign result.
Input: df events fixture Output: user_id | amount 1 | 150.0 2 | 220.0 3 | 110.0 Views, refunds, and purchases <= 100 are dropped.
Topics: lakebench, pyspark, filter, select.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Keep purchase rows with amount > 100 and project user_id, amount.
From `df`, keep rows where `amount > 100` and `event_type == 'purchase'`. Select only `user_id` and `amount`. Order by `user_id`, then `amount`. Assign `result`.