Spark has three primary abstractions. In PySpark, you mostly use DataFrames. Datasets are a Scala/Java concept (typed API); Python does not expose Dataset the same way.
Comparison
RDD (Resilient Distributed Dataset) - typed as Python objects / JVM objects - low-level transformations (map, flatMap, reduceByKey) - little catalyst optimization - use for custom partition logic or legacy code DataFrame - rows with a schema (like a distributed table) - Catalyst + Tungsten optimized - Spark SQL, column expressions, joins, windows - default choice in PySpark Dataset[T] (Scala/Java) - typed records + encoder - compile-time type safety with Catalyst - not the day-to-day PySpark surface
Example
# DataFrame (preferred)
df = spark.read.parquet("/data/events")
df.filter(df.event == "purchase").groupBy("user_id").count()
# RDD (when you truly need it)
rdd = df.rdd.map(lambda row: (row.user_id, 1))
counts = rdd.reduceByKey(lambda a, b: a + b)When to use what
- DataFrame: almost always (ETL, analytics, streaming structured APIs).
- RDD: rare edge cases (custom partitioning, non-relational algorithms, migrating old code).
- Prefer built-in column functions over RDD
mapfor speed and whole-stage codegen.
Common interview answer
"DataFrames sit on RDDs under the hood, but the declarative API lets Catalyst reorder, prune, and codegen the plan. I use DataFrames unless I hit a limitation that only RDDs solve."