PySpark data engineering interview problem. Difficulty: advanced. Pattern: Explain Plans. About 20 minutes. Part of the Pro drill bank.
Cache a frame reused by two aggregations and union the summaries. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Cache df, compute region totals and date totals, union them with a grain column (region|date). For date rows, place the date string in the region column slot. Order by grain, region. Assign result.
Input: reuse df Output: grain | region | total date | 2024-01-01 | 30.0 date | 2024-01-02 | 70.0 date | 2024-01-03 | 110.0 region | APAC | 60.0 region | EU | 60.0 region | US | 90.0 Two actions share a cached base.
Topics: lakebench, pyspark, cache, reuse.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Cache a frame reused by two aggregations and union the summaries.
Cache `df`, compute region totals and date totals, union them with a `grain` column (`region`|`date`). For date rows, place the date string in the `region` column slot. Order by `grain`, `region`. Assign `result`.