PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Aggregation. About 16 minutes. Part of the Pro drill bank.
Active users by signup cohort and transaction month. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Join users to txns. Cohort = first 7 chars of signup_date (YYYY-MM). txn_month similarly from transaction_date. Return cohort, txn_month, active_users (distinct users with a purchase that month). Order by both keys. Assign result.
Input: signups + purchases Output: cohort | txn_month | active_users 2024-01 | 2024-02 | 2 2024-01 | 2024-03 | 1 2024-02 | 2024-02 | 1 2024-02 | 2024-03 | 1 January cohort has two actives in February.
Topics: lakebench, pyspark, cohort, countDistinct.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Active users by signup cohort and transaction month.
Join `users` to `txns`. Cohort = first 7 chars of `signup_date` (YYYY-MM). txn_month similarly from `transaction_date`. Return `cohort`, `txn_month`, `active_users` (distinct users with a purchase that month). Order by both keys. Assign `result`.