Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Learn
  3. Pandas for Data Manipulation

Learn · pandas

Pandas for Data Manipulation

Table transformations in Python, for data that fits on one machine.

DataFrames from scratch: reading data, filtering, grouping, merging, reshaping, and a messy-data silver capstone. SQL comparisons at every step.

20 lessons5 modules5 stages3h 34m
Start lesson 1What is pandas?

Core foundational modules are free. Advanced production modules need Pro.

Why this track exists

You have a 200 MB export with inconsistent country codes, three date formats, and duplicate customer rows. Writing nested loops over that in plain Python takes fifty lines and runs slowly. The work is really table shaped: filter these rows, group by that column, join on this key. Pandas gives you the table operations of SQL inside Python, which is why most real pipeline transformation code looks like this.

Pandas is what you use for the awkward middle: too messy for SQL, too small to justify a Spark cluster. It is also the default tool for investigating a data quality complaint, because you can load a sample and poke at it in minutes.

What you need before starting

  • Python basics

    You need to be comfortable with variables, lists, dicts, and functions. If not, do the first three modules of Core Python first.

  • Helpful: any SQL

    Every operation here is compared to its SQL equivalent, so knowing SQL makes this track much faster. Not required.

The roadmap

5 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.

01The DataFrame itself

Almost every pandas error a beginner hits comes from not knowing what a Series is, what the index is doing, or when they are looking at a copy instead of the original.

Getting Started with PandasFree

0/4

What pandas is, when to use it versus SQL, how to load data, and when apply is the wrong tool.

  1. What is pandas?10m
  2. Pandas versus SQL10m
  3. Reading CSV, Excel, and JSON10m
  4. apply and map12m

By the end of this stage

You can load a file into a DataFrame, select rows and columns deliberately, and predict what an operation returns.

Where people get stuck

The index is the thing people ignore and then fight for weeks. Pay attention to it early and later joins and groupbys stop being mysterious.

02Cleaning real data

Real files have missing values, wrong types, duplicate rows, and inconsistent strings. This stage is the daily work of the job.

Frames & cleaningFree

0/3

Series, DataFrames, missing values, and filters, all on df_orders / df_events.

  1. Series & DataFrames8m
  2. Data cleaning & type casting10m
  3. Filtering & slicing8m

By the end of this stage

You can take a messy export and produce a typed, deduplicated frame, and say what you did to every bad row rather than dropping them silently.

03Reshaping and combining

Data almost never arrives in the shape the question needs. Grouping, merging, pivoting, and long versus wide are the transformations that get it there.

Transforms & shape

0/6

Derived columns, groupby, joins, pivots, and time series.

  1. Transformations & derived columns10m
  2. Grouping & aggregations10m
  3. Combining datasets10m
  4. Reshaping & pivot tables10m
  5. Time series in pandas10m
  6. Memory optimization10m

By the end of this stage

You can group and aggregate, merge frames without silently multiplying rows, and reshape between long and wide on purpose.

Where people get stuck

A merge that duplicates rows is the same bug as a bad SQL join, and just as easy to miss. Check the shape before and after.

04When pandas starts to hurt

Pandas holds everything in memory. There is a size where your laptop stops coping, and knowing where that line is stops you from either wasting a cluster or crashing a job.

Scaling Up

0/3

Parquet and chunks, advanced groupby, and MultiIndex, still on a laptop-sized warehouse.

  1. Parquet, chunks, and out-of-core12m
  2. Advanced groupby: transform, filter, custom agg12m
  3. MultiIndex / hierarchical indexing12m

By the end of this stage

You can reduce memory with dtypes and chunking, and say clearly when a dataset has outgrown pandas and needs Spark.

05Production pandas

Notebook pandas and pipeline pandas are different disciplines. One is for exploring, the other has to run unattended and produce the same answer every time.

Production Pandas

0/4

Schema contracts, SQL round-trip thinking, text cleaning, and a messy-data capstone.

  1. Data validation & schema contracts12m
  2. SQL ⟷ pandas round-trip10m
  3. The .str accessor for text cleaning12m
  4. Capstone: messy data to analysis-ready16m

By the end of this stage

You finish with a silver layer transformation that handles messy input reproducibly, which is exactly the shape of real transformation code.

How you know it worked

Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.

  • You can explain the difference between a Series and a DataFrame without hesitating.
  • You can clean a messy CSV and account for every row you dropped.
  • You can do a groupby with multiple aggregations from memory.
  • You can say, with a number attached, when a dataset is too big for pandas.

How long it takes

30 minutes a day

about 8 sessions

1 hour a day

about 4 sessions

4 hours a weekend day

about 1 session

Keep a scratch dataset open and try each operation on it as you read. Pandas has many ways to do the same thing, and the only way to build judgement about which to reach for is to try them and see what breaks.

These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.

What interviewers are really testing

  • Whether you check the shape after a merge.
  • Whether you know why chained indexing gives you warnings, and what a copy versus a view means.
  • Whether you can translate between SQL and pandas fluently, because interviewers often ask for both.
  • Whether you know when to stop using pandas. Saying 'this needs Spark, here is why' is a senior signal.

Mistakes to avoid on this track

Common mistakes on this track and what to do instead
Common mistakeWhat to do instead
Using pandas for something SQL should do in the warehouse.If the data is already in the warehouse and the operation is a join or an aggregate, do it there. Pulling it out to a laptop is slower and harder to schedule.
Dropping nulls with dropna() without checking what was lost.Count first. Silently deleting 30 percent of your rows is a data incident, not a cleaning step.
Writing loops over rows with iterrows.Vectorised operations are both faster and clearer. If you are looping, there is usually a column operation you have not found yet.

Where to practise this

Pandas drills

Filtered to pandas problems with checked output.

Production tickets

Messy real-world frames with a bug to find.

Where to go after this

PySpark for Distributed Processing

The same DataFrame ideas, distributed across machines, for when the data no longer fits.

Data Quality & Observability

You now know how to clean data. Next: how to prove it stayed clean.

All tracksFull data engineering roadmap45-day plan

Pandas for Data Manipulation reviews & rating

4.9out of 5
1,240+ student reviews
5 stars
88%
4 stars
9%
3 stars
2%
2 stars
1%
1 star
0%
LakeBenchPractice today. Build tomorrow.

Warehouse practice that runs in the tab, not on a cluster. Learn concepts, solve interview drills, and mock the round in one place.

Product

  • Studio sandbox
  • Capstone projects

Practice

  • SQL interview questions
  • PySpark interview questions
  • Python interview questions
  • DE theory questions
  • LeetCode for data engineers

Company

  • About
  • Contact

Legal

  • Privacy
  • Terms
  • Refunds & cancellation
  • Shipping & delivery

© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.

No cluster. No install. Just the tab.