A Pandas UDF (vectorized UDF) applies Python logic to Arrow-backed batches of data as pandas.Series / DataFrame, instead of one row at a time. It is usually much faster than classic Python UDFs.
Example (Series to Series)
import pandas as pd
from pyspark.sql.functions import pandas_udf
@pandas_udf("double")
def add_tax(s: pd.Series) -> pd.Series:
return s * 1.18
df.withColumn("amount_taxed", add_tax("amount"))Types (conceptual)
Series -> Series map-style column transform Series -> Scalar aggregation-style Iterator variants streaming batches for large data Grouped map/agg split-apply-combine with pandas
Diagram
Spark column batches --Arrow--> pandas.Series --> your vectorized code --> Arrow --> Spark
Caveats
- Still Python-side: slower than pure JVM built-ins
- Watch memory: large batches + pandas copies
- Type hints / return types must match Spark schema
- Some operations still better as native Spark expressions
Interview tip
Contrast classic UDF (row) vs Pandas UDF (batch) vs native functions (best).