PySpark data engineering interview problem. Difficulty: beginner. Pattern: Pivot. About 12 minutes. Free to practice.
Widen product amounts per date with conditional aggregation (no DF.pivot). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
The shim has no real pivot. From long df (date, product, amount), build one row per date with columns A, B, C equal to the sum of amounts for that product (0 if missing). Order by date. Assign result.
Input: long product sales Output: date | A | B | C 2024-01-01 | 10.0 | 5.0 | 0.0 2024-01-02 | 7.0 | 0.0 | 3.0 2024-01-03 | 0.0 | 8.0 | 2.0 Missing product/date pairs become 0 via otherwise(0).
Topics: lakebench, pyspark, conditional aggregation, when.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Widen product amounts per date with conditional aggregation (no DF.pivot).
The shim has no real `pivot`. From long `df` (`date`, `product`, `amount`), build one row per `date` with columns `A`, `B`, `C` equal to the sum of amounts for that product (0 if missing). Order by `date`. Assign `result`.