Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. RDD vs DataFrame vs Dataset

Batch & Streaming · Batch Processing

RDD vs DataFrame vs Dataset

Easystream-13
RDDDataFrameDatasetSpark

Question

What are RDDs, DataFrames, and Datasets in Spark, and when would you use each?

Solution

RDD (Resilient Distributed Dataset): low-level distributed collection of JVM objects with lineage. You control partitions and functional transforms. Little automatic optimization.

DataFrame: a distributed table with named columns and a schema. Spark's Catalyst optimizer plans the query. Most DE work lives here (PySpark DataFrame).

Dataset: typed DataFrame-like API on the JVM (Scala/Java). Combines compile-time types with Catalyst. In Python you basically stick to DataFrames.

RDD        : [RowObj, RowObj, ...]   manual, flexible, less optimized
DataFrame  : | id | city | amt |     schema + SQL optimizer
Dataset[T] : typed records (JVM)     types + optimizer

When to choose

  • Default: DataFrame / Spark SQL
  • RDD: rare custom partitioning, legacy code, or APIs not exposed higher up
  • Dataset: Scala/Java typed pipelines

Interview tip: "Prefer DataFrames for optimization and SQL. RDDs are the older foundation."

PreviousNext