RDD (Resilient Distributed Dataset): low-level distributed collection of JVM objects with lineage. You control partitions and functional transforms. Little automatic optimization.
DataFrame: a distributed table with named columns and a schema. Spark's Catalyst optimizer plans the query. Most DE work lives here (PySpark DataFrame).
Dataset: typed DataFrame-like API on the JVM (Scala/Java). Combines compile-time types with Catalyst. In Python you basically stick to DataFrames.
RDD : [RowObj, RowObj, ...] manual, flexible, less optimized DataFrame : | id | city | amt | schema + SQL optimizer Dataset[T] : typed records (JVM) types + optimizer
When to choose
- Default: DataFrame / Spark SQL
- RDD: rare custom partitioning, legacy code, or APIs not exposed higher up
- Dataset: Scala/Java typed pipelines
Interview tip: "Prefer DataFrames for optimization and SQL. RDDs are the older foundation."