Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Built-in functions before UDFs

PySpark · DataFrame API in Practice

Built-in functions before UDFs

Mediumpyspark-79
udfpandas-udfcatalystperformance

Question

Why do senior engineers push back when they see a Python UDF in a PR?

Solution

A Python UDF runs your function in a Python process, one row at a time, outside Spark's optimized engine. That is slow, and it hides logic from the optimizer. So reviewers ask whether a built-in function can do the same thing.

What a UDF costs

  • Serialization: every row is sent from the JVM to a Python worker and back (see the question about how PySpark talks to the JVM).
  • No codegen. The JVM can fuse built-in operators into one fast loop, but the UDF breaks that chain.
  • A black box for Catalyst. Spark cannot push a filter through it, reorder it, or prune columns based on what it needs, so more data may be read and carried.
  • Surprising NULL handling. A Python function that expects a number gets None and raises an exception, which fails the task.

A typical swap

# UDF version
@F.udf("string")
def domain(email):
    return email.split("@")[1].lower() if email else None

# built-in version, same result, no Python
F.lower(F.split("email", "@").getItem(1))

The second version stays in the JVM, and is often several times faster on large data.

What to try before a UDF

  • Functions in pyspark.sql.functions: when, regexp_extract, split, concat_ws, to_date, array_*, and many more.
  • F.expr() with SQL expressions for things the Python API does not wrap.
  • Higher-order functions (transform, filter, aggregate) for arrays.

When a UDF is justified

When the logic really needs a Python library, such as a machine learning model or a geocoding routine. Then use a Pandas UDF (@pandas_udf), which processes a batch of rows as a pandas Series through Apache Arrow. It cuts the per-row overhead a lot.

What to say

Mention measuring. Run the job both ways on realistic data and compare stage time in the UI, rather than assuming. And mention testing: a built-in expression is easier to reason about than a function with hidden exceptions.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext