Snowpark is a set of libraries that lets you write data processing in Python, Java or Scala, and run it inside Snowflake instead of pulling data out to your own machine or a Spark cluster. The DataFrame code you write is translated into SQL and executed by Snowflake's engine.
DataFrame style
from snowflake.snowpark import Session
from snowflake.snowpark.functions import col, sum as sum_
session = Session.builder.configs(conn_params).create()
df = (session.table("orders")
.filter(col("order_date") >= "2025-01-01")
.group_by("region")
.agg(sum_("amount").alias("revenue")))
df.write.save_as_table("region_revenue", mode="overwrite")Nothing runs locally. Like Spark, the operations are lazy, and the work happens in a warehouse when you call an action or write the result. The data never leaves Snowflake.
UDFs and stored procedures
You can also write functions and procedures in Python that run inside Snowflake's secure sandbox, next to the data, with access to many common packages through Snowflake's Anaconda channel. Use a Python UDF for row-level logic that SQL cannot express, and a stored procedure for multi-step workflows.
Snowpark-optimized warehouses
Standard warehouses have limited memory per node. Snowpark-optimized warehouses have much more memory per node, which suits memory-hungry work such as training a model or large in-memory pandas operations.
Why use it
- No data extraction. You do not export tables to run Python elsewhere, so you avoid copying sensitive data and avoid the transfer time.
- One platform and one governance model, since security rules and roles still apply.
- Familiar DataFrame code for Python teams.
When it is not the answer
If you already have large Spark pipelines, rewriting them has a cost. Logic that is easy in SQL should stay in SQL, since plain SQL is often faster and cheaper than a Python UDF row by row. As with Spark, prefer built-in functions to custom UDFs, and measure.