Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a Pandas UDF?

PySpark · UDFs & Streaming

What is a Pandas UDF?

Hardpyspark-32
pandas-udfarrowudfperformance

Question

What is a Pandas UDF (vectorized UDF) in PySpark?

Solution

A Pandas UDF (vectorized UDF) applies Python logic to Arrow-backed batches of data as pandas.Series / DataFrame, instead of one row at a time. It is usually much faster than classic Python UDFs.

Example (Series to Series)

import pandas as pd
from pyspark.sql.functions import pandas_udf

@pandas_udf("double")
def add_tax(s: pd.Series) -> pd.Series:
    return s * 1.18

df.withColumn("amount_taxed", add_tax("amount"))

Types (conceptual)

Series -> Series      map-style column transform
Series -> Scalar      aggregation-style
Iterator variants     streaming batches for large data
Grouped map/agg       split-apply-combine with pandas

Diagram

Spark column batches --Arrow--> pandas.Series --> your vectorized code --> Arrow --> Spark

Caveats

  • Still Python-side: slower than pure JVM built-ins
  • Watch memory: large batches + pandas copies
  • Type hints / return types must match Spark schema
  • Some operations still better as native Spark expressions

Interview tip

Contrast classic UDF (row) vs Pandas UDF (batch) vs native functions (best).

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext