Write the transformation as a function that takes data and returns data, with no hidden reads or writes. Then a test is just: build a small input, call the function, and compare with the output you expect.
A function that is easy to test
import pandas as pd
def add_order_total(df: pd.DataFrame) -> pd.DataFrame:
out = df.copy()
out["total"] = out["qty"] * out["unit_price"]
return outThe test
import pandas as pd
from pandas.testing import assert_frame_equal
from etl.transforms import add_order_total
def test_add_order_total():
given = pd.DataFrame({"qty": [2, 1], "unit_price": [5.0, 3.0]})
expected = pd.DataFrame({"qty": [2, 1], "unit_price": [5.0, 3.0], "total": [10.0, 3.0]})
assert_frame_equal(add_order_total(given), expected)assert_frame_equal checks values, column order and dtypes, and prints a clear difference when it fails. For Spark, the equivalent helpers are pyspark.testing.assertDataFrameEqual or the chispa library.
Cover the awkward cases
Use pytest.mark.parametrize to run the same test over many inputs, so edge cases are cheap to add:
@pytest.mark.parametrize("qty, price, total", [(0, 5.0, 0.0), (3, 0.0, 0.0), (None, 5.0, None)])
def test_totals(qty, price, total): ...Think about what breaks in production: empty input, NULLs, duplicate keys, negative numbers, strings with spaces, timestamps around midnight and daylight saving changes, and an unexpected extra column.
Fixtures and mocks
Use pytest fixtures for shared setup, such as a small sample DataFrame or a temporary directory (tmp_path). Keep real I/O at the edges. Mock or replace the database, API and storage calls (unittest.mock, or a small fake class), so the test checks your logic and not the network.
What to test
Test behaviour that matters: the business rules, joins that could multiply rows, and a bug you fixed once (add its test so it cannot return). Do not test the library itself. Tests should be fast, so developers run them on every change, and CI should run them on every pull request.