agg collapses each group to a single row. transform returns a result with the same number of rows as the input, with each row getting a value computed from its group. If you know SQL, agg is GROUP BY, and transform is a window function with PARTITION BY.
Example data
category product revenue A p1 100 A p2 300 B p3 200
agg: one row per group
df.groupby("category")["revenue"].agg("sum")
# A 400
# B 200transform: same shape as the input
df["category_total"] = df.groupby("category")["revenue"].transform("sum")
df["share"] = df["revenue"] / df["category_total"]Result:
category product revenue category_total share A p1 100 400 0.25 A p2 300 400 0.75 B p3 200 200 1.00
Each row keeps its own identity and gains the group's total. That makes "share of the category", "difference from the group average" and "rank within the group" simple. The SQL equivalent is SUM(revenue) OVER (PARTITION BY category).
More examples
df["z"] = df.groupby("category")["revenue"].transform(lambda s: (s - s.mean()) / s.std())
df["rank"] = df.groupby("category")["revenue"].rank(ascending=False)
df["prev"] = df.groupby("category")["revenue"].shift(1) # like LAG
df["running"] = df.groupby("category")["revenue"].cumsum() # running totalMany of these have their own methods (rank, shift, cumsum) that behave like transform, and are faster than passing a custom lambda.
Performance tip
A string name such as transform("sum") uses an optimised path. A Python lambda is called per group and is much slower when you have many groups.
The alternative to agg then merge
Without transform, you would aggregate, then merge the result back to the original table on the group key. transform does that in one step and avoids a join that could multiply rows.