Overview
pandas lives on one machine. Spark splits work across a cluster. Most tables still fit in pandas; Spark is for when they do not.
On this page7 sections
The decision
The previous lesson walked the full architecture: driver, executors, partitions, lazy plans, shuffle, and stages. This lesson is the short hands-on version of the same idea. Spark is a tool for processing data that does not fit comfortably on one computer. pandas loads a table into your laptop's memory and works there. Spark splits the table into chunks, sends those chunks to many machines, and combines the answers. You still write table verbs (filter, select, count). The work can happen across a cluster.
Most datasets you will touch this week fit in pandas. A few million rows of orders is not a Spark problem. Spark starts to matter when the files are tens of gigabytes, when a join blows up memory, or when the job must finish overnight on a schedule that a laptop cannot keep.
This track teaches the Spark DataFrame API. This tab simulates that API. There is no live cluster. You practice the verbs and the decisions. On a real cluster, the same calls read from object storage and run on executors.
What is at stake
An online store can grow from a spreadsheet of daily orders into years of click events. pandas will load last week without complaint. A year of events can exceed RAM, freeze the kernel, or take so long that the laptop is useless for anything else. Teams then copy the same logic to Spark so many machines share the scan.
Choosing Spark too early has a cost too. A cluster needs setup, billing, and a different debugging story. If the table fits in memory and the job is a one-off, pandas is the faster path from question to answer. A data engineer needs both: pandas for small and local, Spark when one machine is not enough.
Option A vs Option B
Think of pandas as one person with one desk. Spark is a warehouse with many rooms. The desk is fine for a binder. The warehouse is for pallets.
Same table verbs. Different place the data lives while you compute.
Memory is the usual limit. A 2 GB CSV can become a much larger DataFrame after parsing. If your machine has 16 GB of RAM, a few copies of that frame plus joins will crash. Spark keeps partitions on many executors so no single process must hold everything.
Start small. Move to Spark when size, runtime, or shared production files demand it.
| Situation | Use pandas | Use Spark |
|---|---|---|
| Fits in laptop RAM with room to spare | Yes | Usually no need |
| Repeated production job on growing files | Until it does not fit or is too slow | Yes, on a cluster |
| Ad-hoc exploration of a sample | Yes | Overkill |
| Joins and aggregations on a lake of Parquet | Sample first | Yes, when the full grain is required |
The orders table here is tiny on purpose. You still load it with spark.table, the same name you will use later.
A worked comparison
Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.
df = spark.table("orders")
print("row count", df.count())
result = df.limit(5)
result.show()spark.table("orders") is how this tab hands you a DataFrame. count() asks how many rows. limit(5) bounds what you send to the result grid. On a cluster, count() would scan partitions across machines. Here it counts a local sample so you can learn the call.
# pandas, for comparison (not the exercise):
# import pandas as pd
# pdf = pd.read_parquet("orders.parquet")
# print(len(pdf))
# pdf.head(5)
df = spark.table("orders")
print("Spark path in this tab: spark.table, then count and limit")
print("count =", df.count())
result = df.limit(5)
result.show()pandas uses len(df) and head(5). Spark uses count() and limit(5) plus show(). You will see that translation in every lesson.
This tab is not a cluster
Nothing here proves your laptop can handle a terabyte. Practice the API. Judge real size with file size, row counts, and a cluster job, not with this sample.
If you know Pandas, here is the translation
pd.read_parquet(...) plus len(df) plus df.head(5) becomes spark.table("orders") (or spark.read.parquet on a cluster), df.count(), and df.limit(5).show(). pandas holds the frame in RAM. Spark holds a plan until an action.
Before you start
This track assumes you completed Core Python. The pandas track is strongly recommended so the translations land. You do not need a cloud account to complete these exercises.
Copy-paste without reading the output
Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.
Common beginner questions
Can I just use Excel?
Spreadsheets are fine for small, human-edited tables. They do not version well, they choke on millions of rows, and they are not how production pipelines run. pandas is the next step on one machine. Spark is the step after that.
How do I know the data is too big?
Watch RAM, run time, and whether you are sampling because the full table will not load. A 100 MB Parquet file is usually pandas. Many files that add up to hundreds of GB are Spark.
Why learn Spark if this tab is small?
Production data is not small. The API, laziness, and "do not collect everything" habits are what you take to a cluster. The sample keeps feedback fast.
What comes next
If Spark runs on many machines, you need names for those machines. The next lesson is the cluster: driver, executors, and partitions, in beginner language.
Practice
Run Sample to load orders. Then complete Exercise: load spark.table("orders"), print the count, and set result to a 5-row limit of that DataFrame.
You are proving you can open a Spark table and bound the sample. No filter is required yet.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.