Learn · pyspark
Table transformations when the data is too big for one machine.
When data gets too big for one machine: Spark architecture, DataFrames, joins, the medallion pattern, and a bronze-to-gold capstone. Pandas translations included.
Core foundational modules are free. Advanced production modules need Pro.
Playable walkthroughs: watch the system move, predict the next step, stamp a memory seal, then practice. Completing a Trace counts toward readiness.
Your pandas job has been fine for a year. The company grows, the daily file becomes 400 GB, and the job dies on a machine with 32 GB of memory. Buying a bigger machine works until it does not. The alternative is to split the data across many machines and have each one work on its slice. That sounds simple until you need a total per customer and the rows for one customer are scattered across forty machines. Spark exists to manage exactly that problem, and most of what feels strange about Spark is a consequence of it.
Spark is the standard tool for large batch processing, and it is where the interesting failures live: jobs that are slow for no obvious reason, one task that never finishes, memory errors that appear only in production. Being the person who can read a Spark UI and say what is wrong is genuinely valuable.
You write PySpark in Python. Comfort with functions and dicts is enough.
Every Spark operation here is compared to its pandas equivalent. Knowing pandas first makes this track much easier.
Helpful: SQL joins and grouping
Spark's hardest performance topics are all about joins and grouping. Knowing what they mean logically lets you focus on the distributed part.
5 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Spark makes no sense as a list of functions. Driver, executor, partition, task, stage, and shuffle are one connected picture, and everything else in the track hangs off it. This is deliberately the zero to one starting point.
PySpark ArchitectureFree
0/5
The starting point: why Spark exists, who executes your code, how data moves, then the first hands-on verbs on a small table.
By the end of this stage
You can draw what happens when you submit a job, and explain why some operations force data to move between machines.
Where people get stuck
Do not skip ahead to the DataFrame API because it looks familiar. Every performance lesson later assumes this picture is in your head.
With the architecture in place, the API is mostly familiar: select, filter, join, group. What is new is that nothing runs until you ask for a result.
Mental model & DataFrames
0/3
Driver vs executors, lazy plans, and the first DataFrame verbs.
Module Checkpoint
1 conceptual questions to verify mastery
By the end of this stage
You can write transformations, understand why lazy evaluation exists, and know which actions trigger real work.
Real data at scale is rarely flat. Event payloads are nested, questions need ranking within groups, and both of those interact with how Spark distributes work.
Aggregations, windows & nested data
0/6
The verbs you will use on every silver job.
Module Checkpoint
3 conceptual questions to verify mastery
By the end of this stage
You can aggregate, use window functions, and flatten nested structures without guessing.
This is the stage the job actually pays for. Shuffle, skew, broadcast joins, partition pruning, caching, and small files are the difference between a 40 minute job and a 4 minute one.
Performance at Scale
0/3
Cache/persist, data skew and salting, and reading a Catalyst-style physical plan.
By the end of this stage
You can take a slow job, find the bottleneck from evidence rather than guesswork, and explain the tradeoff of the fix you chose.
Where people get stuck
Learn these as a connected system, not seven tricks. Skew and shuffle are the same story, and broadcast joins only make sense once you know what a shuffle costs.
Running Spark reliably means partitioned output, sane file sizes, schema changes that do not break yesterday's data, and a bronze to gold layout other people can use.
Production Spark
0/4
UDFs to avoid, Structured Streaming watermarks, Delta MERGE, and an incremental medallion capstone.
By the end of this stage
You finish with a medallion pipeline that goes from raw to serving, with the performance decisions made on purpose.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 10 sessions
1 hour a day
about 5 sessions
4 hours a weekend day
about 2 sessions
Spend real time on the architecture module even though it has no clever code in it. Learners who rush it spend the performance stage memorising rules they cannot apply. After each performance lesson, close the notes and try to explain the mechanism out loud, because that is exactly what the interview asks for.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Treating Spark as pandas with a different import. | The API looks similar, the execution model is not. Ignoring that is why jobs work on samples and fail on real data. |
| Calling cache() everywhere to make things faster. | Caching costs memory. Cache the frame you reuse several times, after the expensive filter, not the raw source. |
| Learning performance tuning as a list of config settings. | Config without a bottleneck is superstition. Find what is actually slow first, then change the one thing that addresses it. |
| Writing output without thinking about file count. | Millions of tiny files make the next job slow and the metadata expensive. Partition and size output deliberately. |