Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a UDF?

PySpark · UDFs & Streaming

What is a UDF?

Hardpyspark-31
udfperformancepython

Question

What is a UDF in PySpark, and what are the downsides?

Solution

A UDF (User Defined Function) is custom code registered so Spark can apply it row-wise (or via SQL) when built-in functions are not enough.

Python UDF example

from pyspark.sql.functions import udf
from pyspark.sql.types import StringType

def bucket(amount):
    if amount is None:
        return "UNK"
    return "HIGH" if amount >= 100 else "LOW"

bucket_udf = udf(bucket, StringType())
df.withColumn("bucket", bucket_udf("amount"))

Why they hurt performance

Catalyst/Tungsten optimized ops   vs   Python UDF
  whole-stage codegen                    row-by-row
  JVM execution                          Python worker process + serialization

Python UDFs break many optimizations, add serialization between JVM and Python, and prevent whole-stage codegen from fusing operators.

Prefer

  • Built-in pyspark.sql.functions
  • SQL expressions
  • when/otherwise, higher-order functions on arrays
  • Pandas UDFs / Spark Connect vectorized paths when custom logic is required

Interview tip

"UDFs are an escape hatch. I reach for them last and prefer Pandas UDFs or native functions when custom logic is unavoidable."

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext