Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. apply vs vectorization

Python · pandas & Polars

apply vs vectorization

Mediumpython-54
pandasapplyvectorizationperformancenumpy

Question

Why is df.apply(row_function, axis=1) slow, and what should you use instead?

Solution

df.apply(func, axis=1) calls your Python function once for every row, and Python function calls are slow. A vectorised operation hands the whole column to compiled code (NumPy, written in C), which loops over the data much faster than Python can.

The slow way and the fast way

# slow: a Python function call per row
df["total"] = df.apply(lambda r: r["qty"] * r["unit_price"], axis=1)

# fast: one operation on two whole columns
df["total"] = df["qty"] * df["unit_price"]

On 10 million rows, the first can take tens of seconds, and the second about a tenth of a second. The gap is typically one to two orders of magnitude, and it grows with the data. Measure on your own machine with %timeit before and after.

Why apply is slow

For each row, pandas builds a Series object, calls your function, and collects the result. That is a lot of interpreter overhead, and none of it can use the CPU's fast array instructions. axis=1 is the worst case.

What to use instead

  • Column arithmetic and comparisons: df["a"] + df["b"], df["x"] > 0.
  • np.where or np.select for if/else logic:
df["tier"] = np.where(df["amount"] > 1000, "high", "low")

More vectorised tools:

  • The .str and .dt accessors for text and dates: df["email"].str.lower(), df["ts"].dt.date.
  • merge for lookups, instead of a function that searches another table for every row.
  • map with a dictionary for simple code-to-label conversion: df["code"].map(mapping).
  • groupby().transform() for group-level calculations.

When apply is acceptable

When the logic truly cannot be expressed with column operations (complex parsing of a text blob, a call to a library that works on one item), and the data is small. Even then, try apply on a single column (df["col"].apply(f)), which is cheaper than row-wise.

Bigger picture

If you find yourself writing per-row Python logic on large data, that is the cue to use Polars expressions, DuckDB SQL or Spark, which keep the work in compiled code. Say that the habit to build is thinking in columns, not rows.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext