PySpark is the Python API for Apache Spark, a distributed data processing engine. You write Python code; Spark runs the heavy work across many machines (or many cores on one machine) as parallel tasks.
Why people use it
Python is popular for data work, but a single Python process cannot scan terabytes efficiently. PySpark lets you keep Python syntax while Spark handles partitioning, scheduling, shuffle, and fault recovery.
Mental model
Your Python code (driver)
|
v
SparkSession / JVM
|
+--> tasks on executor 1
+--> tasks on executor 2
+--> tasks on executor NMinimal example
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("demo").getOrCreate()
df = spark.read.parquet("s3://bucket/orders/")
df.groupBy("region").count().show()Interview talking points
- PySpark is not "Python Spark from scratch"; it talks to the Spark JVM via Py4J.
- Prefer DataFrame / Spark SQL APIs over Python UDFs for performance.
- Same engine powers batch ETL, interactive analytics, and Structured Streaming.
Common confusion
PySpark is not the same as pandas. pandas loads data into one machine's memory. PySpark plans distributed execution and can spill / shuffle across the cluster.