Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Snowpark

Snowflake, BigQuery & Databricks · Snowflake

Snowpark

Mediumwarehouses-18
snowflakesnowparkpythonudfstored-procedures

Question

What is Snowpark?

Solution

Snowpark is a set of libraries that lets you write data processing in Python, Java or Scala, and run it inside Snowflake instead of pulling data out to your own machine or a Spark cluster. The DataFrame code you write is translated into SQL and executed by Snowflake's engine.

DataFrame style

from snowflake.snowpark import Session
from snowflake.snowpark.functions import col, sum as sum_

session = Session.builder.configs(conn_params).create()

df = (session.table("orders")
      .filter(col("order_date") >= "2025-01-01")
      .group_by("region")
      .agg(sum_("amount").alias("revenue")))

df.write.save_as_table("region_revenue", mode="overwrite")

Nothing runs locally. Like Spark, the operations are lazy, and the work happens in a warehouse when you call an action or write the result. The data never leaves Snowflake.

UDFs and stored procedures

You can also write functions and procedures in Python that run inside Snowflake's secure sandbox, next to the data, with access to many common packages through Snowflake's Anaconda channel. Use a Python UDF for row-level logic that SQL cannot express, and a stored procedure for multi-step workflows.

Snowpark-optimized warehouses

Standard warehouses have limited memory per node. Snowpark-optimized warehouses have much more memory per node, which suits memory-hungry work such as training a model or large in-memory pandas operations.

Why use it

  • No data extraction. You do not export tables to run Python elsewhere, so you avoid copying sensitive data and avoid the transfer time.
  • One platform and one governance model, since security rules and roles still apply.
  • Familiar DataFrame code for Python teams.

When it is not the answer

If you already have large Spark pipelines, rewriting them has a cost. Logic that is easy in SQL should stay in SQL, since plain SQL is often faster and cheaper than a Python UDF row by row. As with Spark, prefer built-in functions to custom UDFs, and measure.

PreviousNext