A UDF (User Defined Function) is custom code registered so Spark can apply it row-wise (or via SQL) when built-in functions are not enough.
Python UDF example
from pyspark.sql.functions import udf
from pyspark.sql.types import StringType
def bucket(amount):
if amount is None:
return "UNK"
return "HIGH" if amount >= 100 else "LOW"
bucket_udf = udf(bucket, StringType())
df.withColumn("bucket", bucket_udf("amount"))Why they hurt performance
Catalyst/Tungsten optimized ops vs Python UDF whole-stage codegen row-by-row JVM execution Python worker process + serialization
Python UDFs break many optimizations, add serialization between JVM and Python, and prevent whole-stage codegen from fusing operators.
Prefer
- Built-in
pyspark.sql.functions - SQL expressions
when/otherwise, higher-order functions on arrays- Pandas UDFs / Spark Connect vectorized paths when custom logic is required
Interview tip
"UDFs are an escape hatch. I reach for them last and prefer Pandas UDFs or native functions when custom logic is unavoidable."