Learn · pandas
Table transformations in Python, for data that fits on one machine.
DataFrames from scratch: reading data, filtering, grouping, merging, reshaping, and a messy-data silver capstone. SQL comparisons at every step.
Core foundational modules are free. Advanced production modules need Pro.
You have a 200 MB export with inconsistent country codes, three date formats, and duplicate customer rows. Writing nested loops over that in plain Python takes fifty lines and runs slowly. The work is really table shaped: filter these rows, group by that column, join on this key. Pandas gives you the table operations of SQL inside Python, which is why most real pipeline transformation code looks like this.
Pandas is what you use for the awkward middle: too messy for SQL, too small to justify a Spark cluster. It is also the default tool for investigating a data quality complaint, because you can load a sample and poke at it in minutes.
You need to be comfortable with variables, lists, dicts, and functions. If not, do the first three modules of Core Python first.
Every operation here is compared to its SQL equivalent, so knowing SQL makes this track much faster. Not required.
5 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Almost every pandas error a beginner hits comes from not knowing what a Series is, what the index is doing, or when they are looking at a copy instead of the original.
Getting Started with PandasFree
0/4
What pandas is, when to use it versus SQL, how to load data, and when apply is the wrong tool.
By the end of this stage
You can load a file into a DataFrame, select rows and columns deliberately, and predict what an operation returns.
Where people get stuck
The index is the thing people ignore and then fight for weeks. Pay attention to it early and later joins and groupbys stop being mysterious.
Real files have missing values, wrong types, duplicate rows, and inconsistent strings. This stage is the daily work of the job.
Frames & cleaningFree
0/3
Series, DataFrames, missing values, and filters, all on df_orders / df_events.
By the end of this stage
You can take a messy export and produce a typed, deduplicated frame, and say what you did to every bad row rather than dropping them silently.
Data almost never arrives in the shape the question needs. Grouping, merging, pivoting, and long versus wide are the transformations that get it there.
Transforms & shape
0/6
Derived columns, groupby, joins, pivots, and time series.
By the end of this stage
You can group and aggregate, merge frames without silently multiplying rows, and reshape between long and wide on purpose.
Where people get stuck
A merge that duplicates rows is the same bug as a bad SQL join, and just as easy to miss. Check the shape before and after.
Pandas holds everything in memory. There is a size where your laptop stops coping, and knowing where that line is stops you from either wasting a cluster or crashing a job.
Scaling Up
0/3
Parquet and chunks, advanced groupby, and MultiIndex, still on a laptop-sized warehouse.
By the end of this stage
You can reduce memory with dtypes and chunking, and say clearly when a dataset has outgrown pandas and needs Spark.
Notebook pandas and pipeline pandas are different disciplines. One is for exploring, the other has to run unattended and produce the same answer every time.
Production Pandas
0/4
Schema contracts, SQL round-trip thinking, text cleaning, and a messy-data capstone.
By the end of this stage
You finish with a silver layer transformation that handles messy input reproducibly, which is exactly the shape of real transformation code.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 8 sessions
1 hour a day
about 4 sessions
4 hours a weekend day
about 1 session
Keep a scratch dataset open and try each operation on it as you read. Pandas has many ways to do the same thing, and the only way to build judgement about which to reach for is to try them and see what breaks.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Using pandas for something SQL should do in the warehouse. | If the data is already in the warehouse and the operation is a join or an aggregate, do it there. Pulling it out to a laptop is slower and harder to schedule. |
| Dropping nulls with dropna() without checking what was lost. | Count first. Silently deleting 30 percent of your rows is a data incident, not a cleaning step. |
| Writing loops over rows with iterrows. | Vectorised operations are both faster and clearer. If you are looping, there is usually a column operation you have not found yet. |