Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Pandas for Data Manipulation

Progress0/20
x

Getting Started with Pandas

  • What is pandas?10m
  • Pandas versus SQL10m
  • Reading CSV, Excel, and JSON10m
  • apply and map12m

Frames & cleaning

  • Series & DataFrames8m
  • Data cleaning & type casting10m
  • Filtering & slicing8m

Transforms & shape

  • Transformations & derived columns10m
  • Grouping & aggregations10m
  • Combining datasets10m
  • Reshaping & pivot tables10m
  • Time series in pandas10m
  • Memory optimization10m

Scaling Up

  • Parquet, chunks, and out-of-core12m
  • Advanced groupby: transform, filter, custom agg12m
  • MultiIndex / hierarchical indexing12m

Production Pandas

  • Data validation & schema contracts12m
  • SQL ⟷ pandas round-trip10m
  • The .str accessor for text cleaning12m
  • Capstone: messy data to analysis-ready16m
Back to track
  1. Learn
  2. Pandas for Data Manipulation
  3. Getting Started with Pandas
  4. apply and map

Lesson 4 of 20 · Theory first, then run it

apply and map

pandasbeginner12 min

Overview

.map remaps a Series. .apply walks row by row or column by column. Vectorized comparisons are faster than both.

On this page7 sections›
  1. 1The decision
  2. 2What is at stake
  3. 3Option A vs Option B
  4. 4A worked comparison
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The decision

Once a table is in memory, you often need a new column: a flag, a cleaned label, or a mapped code. Pandas gives you three ways to compute that column. Two of them walk values one at a time. One of them hits the whole column at once.

.map lives on a Series. You pass a dictionary or a function. Each value in the column is replaced. The result is still one column.

.apply can run a function down each column (axis=0) or across each row (axis=1). Row-wise apply is a Python loop in disguise. Vectorized work (comparisons, arithmetic, .str methods) does the same job without visiting rows in Python. Prefer the vectorized path.

What is at stake

A common first task is a boolean flag: is this order paid? Beginners reach for apply and a lambda because it looks like a for-loop. On a few hundred rows it seems fine. On a few million rows it is the difference between a second and a coffee break.

Data engineers write the same flag in production notebooks every day. Interviewers also watch for apply used where a comparison would do. Learning the fast path now means later lessons (cleaning with .str, derived columns, groupby) stay vectorized by default.

Option A vs Option B

Slow path versus fast path
Row-wise .apply (Python per row)Series .map (one column, still per value)Vectorized == / .str / arithmetic

Row-wise apply visits Python for every row. A column comparison runs as one operation.

Picture a spray gun versus a toothbrush. Vectorized operations spray the whole column. apply brushes each tile. Use the toothbrush only when no spray exists: a function that truly needs several columns in Python, with no pandas equivalent.

If a comparison or .str method can do it, do not apply.

ToolWorks onTypical useSpeed
.map(dict or function)A SeriesRemap codes: {'paid': 1, 'pending': 0}Fine for one column
.apply(fn, axis=0)Each columnA summary function per columnUsually replaceable
.apply(fn, axis=1)Each rowLogic that needs many columns in PythonSlow; last resort
Vectorized ==, .str, arithmeticA whole columnFlags, math, text cleanupPreferred

Here is the paid flag on a tiny extract. The wrong version calls apply. The right version compares the column to the string paid.

Add is_paid without a row loop
order_idorder_statusis_paidO-104paidTrueO-109pendingFalseO-116paidTrue

is_paid is True only where order_status is paid. Same grain: one row per order.

A worked comparison

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

PythonWrong idea in comments. Right idea: a vectorized comparison
# Slow and unnecessary for this flag (do not turn in this pattern):
# out["is_paid"] = out.apply(lambda row: row["order_status"] == "paid", axis=1)

out = df_orders.copy()
out["is_paid"] = out["order_status"] == "paid"
result = out.loc[:, ["order_id", "order_status", "is_paid"]].head(12)

out["order_status"] == "paid" returns a Series of True/False aligned to the index. Assigning it creates is_paid. No apply, no lambda, no axis=1. Copy first so you do not surprise yourself by editing a slice.

Python.map remaps a Series from a dictionary
status = df_orders["order_status"]
mapped = status.map({"paid": "settled", "refunded": "settled"})
result = df_orders.assign(status_group=mapped).loc[:, ["order_id", "order_status", "status_group"]].head(8)

.map looks up each status in the dictionary. Values that are missing from the dict become null. That is useful for recoding. It is still the wrong tool for a simple equality flag: use == for that.

  1. Copy the frame if you will add columns.
  2. Prefer a vectorized comparison or .str method.
  3. Use .map when you have an explicit lookup table.
  4. Reach for .apply only after you confirm no vectorized method exists.

apply is not the default

axis=1 apply is the first tool many tutorials show and the last tool you should reach for. If you are writing lambda row:, stop and ask whether ==, .isin, .str, or arithmetic already does the job.

SQL connection

CASE WHEN order_status = 'paid' THEN TRUE ELSE FALSE END is the warehouse version of df['order_status'] == 'paid'. SQL engines vectorize that CASE. Pandas vectorizes the comparison. Both are cheaper than a row-by-row loop.

Copy-paste without reading the output

Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.

Common beginner questions

When is apply acceptable?

When the logic has no column-wise equivalent: calling an external API per row, parsing a one-off nested structure, or a function pandas cannot express. Even then, try a vectorized rewrite first.

What does axis mean?

axis=0 means down a column (the default). axis=1 means across a row. For flags like is_paid you should not need axis at all, because you should not need apply.

Is .map faster than apply?

For a dict lookup on one Series, yes, and it is clearer. A boolean flag still should not use either. Use a comparison.

What comes next

You can inspect a frame, compare pandas to SQL, load (or receive) a table, and add a column the fast way. The next module goes deeper on Series, DataFrames, .loc, and .iloc so selection becomes muscle memory.

Practice

Run Sample to see is_paid built with a comparison. Then complete Exercise: copy df_orders, add is_paid with a vectorized check against "paid" (not apply), keep order_id, order_status, and is_paid, take 12 rows, and assign that DataFrame to result.

Checks look for the name is_paid and reject code that contains apply.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Reading CSV, Excel, and JSONSeries & DataFrames