Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. How do you pivot a DataFrame?

PySpark · DataFrame API & I/O

How do you pivot a DataFrame?

Mediumpyspark-17
pivotreshapeaggregation

Question

How do you pivot a DataFrame in PySpark?

Solution

Pivot turns unique values of a column into new columns, usually with an aggregation. Classic use: reshape event counts or monthly metrics into a wide report.

Example

from pyspark.sql import functions as F

sales = spark.createDataFrame(
    [
        ("east", "2024-01", 10),
        ("east", "2024-02", 12),
        ("west", "2024-01", 8),
    ],
    ["region", "month", "amount"],
)

wide = (
    sales.groupBy("region")
    .pivot("month", ["2024-01", "2024-02"])  # optional explicit values
    .sum("amount")
)
wide.show()

Diagram

region month   amount          region  2024-01  2024-02
east   2024-01 10       =>     east    10       12
east   2024-02 12              west    8        null
west   2024-01 8

Tips

  • Pass explicit pivot values when known; discovering distinct values requires an extra pass.
  • Pivot implies aggregation (sum, count, agg).
  • Inverse is stack / melt-style unpivot with expr("stack(...)") or pandas-style reshaping outside Spark.

Interview tip

Wide pivots with high cardinality explode column count and can hurt planning; keep pivot keys bounded.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext