PySpark data engineering interview problem. Difficulty: advanced. Pattern: Explain Plans. About 20 minutes. Part of the Pro drill bank.
Cache df then aggregate totals by region. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Call cache() on df (shim prints a note), then group by region summing amount as total. Order by region. Assign result.
Input: regional amounts Output: region | total APAC | 60.0 EU | 60.0 US | 90.0 Caching does not change aggregate math.
Topics: lakebench, pyspark, cache.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Cache df then aggregate totals by region.
Call `cache()` on `df` (shim prints a note), then group by `region` summing `amount` as `total`. Order by `region`. Assign `result`.