Pivot turns unique values of a column into new columns, usually with an aggregation. Classic use: reshape event counts or monthly metrics into a wide report.
Example
from pyspark.sql import functions as F
sales = spark.createDataFrame(
[
("east", "2024-01", 10),
("east", "2024-02", 12),
("west", "2024-01", 8),
],
["region", "month", "amount"],
)
wide = (
sales.groupBy("region")
.pivot("month", ["2024-01", "2024-02"]) # optional explicit values
.sum("amount")
)
wide.show()Diagram
region month amount region 2024-01 2024-02 east 2024-01 10 => east 10 12 east 2024-02 12 west 8 null west 2024-01 8
Tips
- Pass explicit pivot values when known; discovering distinct values requires an extra pass.
- Pivot implies aggregation (
sum,count,agg). - Inverse is
stack/ melt-style unpivot withexpr("stack(...)")orpandas-style reshaping outside Spark.
Interview tip
Wide pivots with high cardinality explode column count and can hurt planning; keep pivot keys bounded.