PySpark data engineering interview problem. Difficulty: beginner. Pattern: GroupBy. About 12 minutes. Free to practice.
Per-user total, average, count, and max amount. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From df, group by user_id and compute total_amount, avg_amount, txn_count, max_amount. Order by user_id. Assign result.
Input: user product amounts Output: user_id | total_amount | avg_amount | txn_count | max_amount 1 | 60.0 | 20.0 | 3 | 30.0 2 | 20.0 | 10.0 | 2 | 15.0 3 | 8.0 | 8.0 | 1 | 8.0 Multiple agg expressions share one groupBy.
Topics: lakebench, pyspark, groupBy, agg.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Per-user total, average, count, and max amount.
From `df`, group by `user_id` and compute `total_amount`, `avg_amount`, `txn_count`, `max_amount`. Order by `user_id`. Assign `result`.