Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

GroupBy with Multiple Aggregations

PySpark data engineering interview problem. Difficulty: beginner. Pattern: GroupBy. About 12 minutes. Free to practice.

Per-user total, average, count, and max amount. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

From df, group by user_id and compute total_amount, avg_amount, txn_count, max_amount. Order by user_id. Assign result.

Constraints

  • Four named aggregate columns.
  • Order by user_id.

Examples

Input: user product amounts Output: user_id | total_amount | avg_amount | txn_count | max_amount 1 | 60.0 | 20.0 | 3 | 30.0 2 | 20.0 | 10.0 | 2 | 15.0 3 | 8.0 | 8.0 | 1 | 8.0 Multiple agg expressions share one groupBy.

Topics: lakebench, pyspark, groupBy, agg.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

GroupBy with Multiple Aggregations

Interview-style drill: Per-user total, average, count, and max amount.

From `df`, group by `user_id` and compute `total_amount`, `avg_amount`, `txn_count`, `max_amount`. Order by `user_id`. Assign `result`.