Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

When to Cache

PySpark data engineering interview problem. Difficulty: advanced. Pattern: Explain Plans. About 20 minutes. Part of the Pro drill bank.

Cache a frame reused by two aggregations and union the summaries. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Cache df, compute region totals and date totals, union them with a grain column (region|date). For date rows, place the date string in the region column slot. Order by grain, region. Assign result.

Constraints

  • Call cache() once on the shared base.
  • Include both grains.

Examples

Input: reuse df Output: grain | region | total date | 2024-01-01 | 30.0 date | 2024-01-02 | 70.0 date | 2024-01-03 | 110.0 region | APAC | 60.0 region | EU | 60.0 region | US | 90.0 Two actions share a cached base.

Topics: lakebench, pyspark, cache, reuse.

More PySpark interview questions · All interview problems · Learn data engineering

advanced

When to Cache

Interview-style drill: Cache a frame reused by two aggregations and union the summaries.

Cache `df`, compute region totals and date totals, union them with a `grain` column (`region`|`date`). For date rows, place the date string in the `region` column slot. Order by `grain`, `region`. Assign `result`.