SparkSession is the single entry point to Spark functionality in modern Spark (2.x+). It unifies what used to be separate contexts (SparkContext, SQLContext, HiveContext).
What it gives you
- Create DataFrames / Datasets
- Run Spark SQL
- Configure the app (shuffle partitions, AQE, etc.)
- Read/write data sources
- Access the underlying
SparkContextwhen needed (spark.sparkContext)
Create one
from pyspark.sql import SparkSession
spark = (
SparkSession.builder
.appName("orders-etl")
.config("spark.sql.shuffle.partitions", "200")
.config("spark.sql.adaptive.enabled", "true")
.getOrCreate()
)
print(spark.version)
spark.sql("SELECT 1 AS ok").show()Diagram
SparkSession |-- spark.read / spark.write |-- spark.sql(...) |-- spark.catalog / spark.udf +-- spark.sparkContext (low-level RDD / cluster APIs)
Interview tip
In notebooks and Databricks, a spark session is often pre-created. In production jobs, you still build and configure it explicitly so settings are versioned with the code.