Write your transformations as plain functions that take a DataFrame and return a DataFrame, create a small local SparkSession once for the whole test run, build tiny input DataFrames in the test, and compare the result with an expected DataFrame.
Structure the code for testing
# transforms.py
from pyspark.sql import DataFrame, functions as F
def add_order_total(df: DataFrame) -> DataFrame:
return df.withColumn("total", F.col("qty") * F.col("unit_price"))A function like this does not read from or write to anywhere, so the test needs no files or cluster. Keep I/O (reading and writing) in a thin outer layer.
A pytest fixture
# conftest.py
import pytest
from pyspark.sql import SparkSession
@pytest.fixture(scope="session")
def spark():
s = (SparkSession.builder
.master("local[2]")
.appName("tests")
.config("spark.sql.shuffle.partitions", "2")
.config("spark.ui.enabled", "false")
.getOrCreate())
yield s
s.stop()Starting Spark takes several seconds, so use scope="session" to pay that once. Lowering the shuffle partitions from 200 to 2 makes small tests much faster.
The test
from pyspark.testing import assertDataFrameEqual
def test_add_order_total(spark):
given = spark.createDataFrame([(2, 5.0), (1, 3.0)], ["qty", "unit_price"])
expected = spark.createDataFrame([(2, 5.0, 10.0), (1, 3.0, 3.0)],
["qty", "unit_price", "total"])
assertDataFrameEqual(add_order_total(given), expected)assertDataFrameEqual is in pyspark.testing from Spark 3.5. It compares schema and rows, ignoring row order by default, and gives a readable diff. For older versions, the chispa library offers similar helpers.
What to cover
Besides the happy path, test the awkward inputs: NULLs, empty DataFrames, duplicate keys, a date at the edge of a month, and an unexpected extra column. Check the schema too, not just the values, since a type change is a real bug. For joins and windows, keep three or four rows that make the expected result easy to check by hand.
In CI
Run the tests on every pull request in a container that has Java and PySpark installed. Keep the data tiny so the suite runs in a minute or two. Heavier checks, such as performance and full-size runs, belong in a separate stage on a real cluster, and not in the unit tests.