Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Cache vs Persist

PySpark data engineering interview problem. Difficulty: advanced. Pattern: Explain Plans. About 20 minutes. Part of the Pro drill bank.

Cache df then aggregate totals by region. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Call cache() on df (shim prints a note), then group by region summing amount as total. Order by region. Assign result.

Constraints

  • Call cache().
  • Sum amount per region.

Examples

Input: regional amounts Output: region | total APAC | 60.0 EU | 60.0 US | 90.0 Caching does not change aggregate math.

Topics: lakebench, pyspark, cache.

More PySpark interview questions · All interview problems · Learn data engineering

advanced

Cache vs Persist

Interview-style drill: Cache df then aggregate totals by region.

Call `cache()` on `df` (shim prints a note), then group by `region` summing `amount` as `total`. Order by `region`. Assign `result`.