Spark Connect
A thin client that talks to a remote Spark server over gRPC, and what it changes for applications.
On this page
You will learn
- What Spark Connect is and how it differs from a classic session
- How a client talks to the server
- What you gain: stability, upgrades, lighter clients
- What does not work over Connect
Classic mode
In a classic PySpark application, your Python process starts a JVM driver locally (through Py4J) and the driver runs inside your application. The client and the driver share a fate and a version: a memory leak in your app can kill the driver, and upgrading Spark means upgrading every application at once.
Connect mode
-
Client
Builds an unresolved plan from DataFrame calls
-
gRPC
Sends the plan (protobuf) to the server
-
Server
A long-running driver analyses, optimises and executes it
-
Arrow
Results stream back as Arrow batches
from pyspark.sql import SparkSession spark = SparkSession.builder.remote("sc://spark-server:15002").getOrCreate() spark.read.table("sales").groupBy("country").count().show()
The server is started with sbin/start-connect-server.sh (default port 15002). Most DataFrame and SQL code runs unchanged; the API is the same pyspark.sql DataFrame API.
Why it matters
Stability
Upgrades
Thin clients
It is also what powers notebooks and IDEs talking to remote clusters, and serverless offerings: the client is just a small library speaking a protocol.
What changes
- No RDD API and no SparkContext on the client:
spark.sparkContext,df.rddand RDD-based libraries do not work. Code must use DataFrames and SQL. - Analysis is deferred: in classic mode a bad column name fails as soon as you write the transformation; over Connect, plans are analysed on the server, typically when an action runs or the schema is requested. Errors can surface later.
- Schema access is a round trip:
df.columnsanddf.schemaask the server. Calling them in a tight loop is slow. - Python UDFs still work: the code is serialised and runs on the server's executors, so the server needs compatible Python and libraries.
Common mistakes
Using df.rdd or sparkContext
Calling df.schema in loops
Assuming errors appear immediately
Key takeaways
- Spark Connect sends unresolved plans over gRPC to a remote driver.
- Results come back as Arrow batches.
- It isolates clients from the driver and decouples upgrades.
- No RDD API or SparkContext on the client; analysis errors can appear later.
Check yourself
3 questions1. What does a Spark Connect client send to the server?
Show the answer
Unresolved logical plans over gRPC. DataFrame calls become protobuf-encoded unresolved plans.
2. Which of these does not work over Spark Connect?
Show the answer
df.rdd.map(...). The RDD API and SparkContext are not available to Connect clients.
3. In which version was Spark Connect introduced?
Show the answer
3.4. Spark Connect arrived in Spark 3.4.
Go deeper
Driver, executors and the cluster managerSpark internals
Python Data Source APISpark internals
Transformations vs actions
Primary sources: Spark Connect overview