Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Spark Connect

PySpark · Streaming & Newer Spark

Spark Connect

Mediumpyspark-89
spark-connectclient-serverarchitecture

Question

What is Spark Connect?

Solution

Spark Connect separates the client from the Spark driver. Your application (a Python script, a notebook, an IDE session) runs as a thin client that sends DataFrame operations over gRPC to a remote Spark server, which runs them and streams results back. It was introduced in Spark 3.4.

The old way and the new way

Classic:   [your Python process = driver + JVM] <--> executors
Connect:   [thin client] --gRPC--> [Spark Connect server (driver)] <--> executors

In the classic model, your code ran in the same process as the driver, with direct access to the JVM. A heavy job or a bug in your code could crash the driver, upgrades required changing the client and cluster together, and a client needed a full Spark installation.

What the client sends

The client does not run Spark code. It builds an unresolved logical plan from your DataFrame calls and sends it to the server, which analyses, optimizes, and executes it. The data comes back in Arrow batches. Because the protocol is language-neutral, clients can exist for Python, Scala, Java, Go, Rust and others, and a small client can be embedded in an application or an IDE.

spark = SparkSession.builder.remote("sc://spark-server:15002").getOrCreate()
spark.range(10).filter("id > 5").show()

Benefits

  • Lighter clients, since no JVM or full Spark distribution is needed on the client side.
  • Better isolation: a crash or out-of-memory in a client does not take down the shared driver, and several clients can share a server.
  • Easier upgrades: the server can move to a newer Spark version while clients stay compatible within the protocol.
  • Easier remote development from a laptop.

What you give up

Anything that depended on the driver being in your process. Direct access to the JVM through Py4J and the SparkContext, many RDD APIs, and some low-level features are not available. Code that uses those has to be rewritten with DataFrame APIs. Also, UDFs and dependencies need to be shipped to the server.

Where you meet it

Databricks Connect and serverless or shared-access compute use this model. Open source Spark offers a Connect server you can start yourself. If an interviewer asks for the point in one sentence: it splits Spark into a thin client and a remote server.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext