PySpark data engineering interview problem. Difficulty: intermediate. Pattern: GroupBy. About 16 minutes. Part of the Pro drill bank.
Users with more than 2 transactions. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Group df by user_id, count rows as txn_count, and keep users with txn_count > 2. Order by user_id. Assign result.
Input: transactions Output: user_id | txn_count 1 | 3 Only user 1 clears the > 2 bar.
Topics: lakebench, pyspark, having, count.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Users with more than 2 transactions.
Group `df` by `user_id`, count rows as `txn_count`, and keep users with `txn_count > 2`. Order by `user_id`. Assign `result`.