Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. RDD vs DataFrame vs Dataset

PySpark · Core Concepts

RDD vs DataFrame vs Dataset

Mediumpyspark-04
rdddataframedatasetapi

Question

Compare RDD, DataFrame, and Dataset in Spark. When would you use each?

Solution

Spark has three primary abstractions. In PySpark, you mostly use DataFrames. Datasets are a Scala/Java concept (typed API); Python does not expose Dataset the same way.

Comparison

RDD (Resilient Distributed Dataset)
  - typed as Python objects / JVM objects
  - low-level transformations (map, flatMap, reduceByKey)
  - little catalyst optimization
  - use for custom partition logic or legacy code

DataFrame
  - rows with a schema (like a distributed table)
  - Catalyst + Tungsten optimized
  - Spark SQL, column expressions, joins, windows
  - default choice in PySpark

Dataset[T] (Scala/Java)
  - typed records + encoder
  - compile-time type safety with Catalyst
  - not the day-to-day PySpark surface

Example

# DataFrame (preferred)
df = spark.read.parquet("/data/events")
df.filter(df.event == "purchase").groupBy("user_id").count()

# RDD (when you truly need it)
rdd = df.rdd.map(lambda row: (row.user_id, 1))
counts = rdd.reduceByKey(lambda a, b: a + b)

When to use what

  • DataFrame: almost always (ETL, analytics, streaming structured APIs).
  • RDD: rare edge cases (custom partitioning, non-relational algorithms, migrating old code).
  • Prefer built-in column functions over RDD map for speed and whole-stage codegen.

Common interview answer

"DataFrames sit on RDDs under the hood, but the declarative API lets Catalyst reorder, prune, and codegen the plan. I use DataFrames unless I hit a limitation that only RDDs solve."

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext