A Python UDF runs your function in a Python process, one row at a time, outside Spark's optimized engine. That is slow, and it hides logic from the optimizer. So reviewers ask whether a built-in function can do the same thing.
What a UDF costs
- Serialization: every row is sent from the JVM to a Python worker and back (see the question about how PySpark talks to the JVM).
- No codegen. The JVM can fuse built-in operators into one fast loop, but the UDF breaks that chain.
- A black box for Catalyst. Spark cannot push a filter through it, reorder it, or prune columns based on what it needs, so more data may be read and carried.
- Surprising NULL handling. A Python function that expects a number gets
Noneand raises an exception, which fails the task.
A typical swap
# UDF version
@F.udf("string")
def domain(email):
return email.split("@")[1].lower() if email else None
# built-in version, same result, no Python
F.lower(F.split("email", "@").getItem(1))The second version stays in the JVM, and is often several times faster on large data.
What to try before a UDF
- Functions in
pyspark.sql.functions:when,regexp_extract,split,concat_ws,to_date,array_*, and many more. F.expr()with SQL expressions for things the Python API does not wrap.- Higher-order functions (
transform,filter,aggregate) for arrays.
When a UDF is justified
When the logic really needs a Python library, such as a machine learning model or a geocoding routine. Then use a Pandas UDF (@pandas_udf), which processes a batch of rows as a pandas Series through Apache Arrow. It cuts the per-row overhead a lot.
What to say
Mention measuring. Run the job both ways on realistic data and compare stage time in the UI, rather than assuming. And mention testing: a built-in expression is easier to reason about than a function with hidden exceptions.