Data engineering teams choose from several testing frameworks beyond Great Expectations, selecting tools that integrate directly with their specific execution environments. Leading alternatives include dbt testing packages, Soda Core, PyDeequ for Apache Spark, Databricks Delta Live Tables expectations, and cloud-native services like Google Cloud Dataplex.
Modern assertion frameworks and engines
Depending on where data transformations execute, different tools offer distinct advantages:
- dbt testing ecosystem: Standard dbt provides built-in assertions for uniqueness, nullability, accepted values, and relationships. Teams extend this with dbt-expectations to run statistical distributions in SQL, dbt-utils for recency and expression tests, and Elementary for automated warehouse anomaly detection.
- Soda Core: An open-source CLI and Python tool using declarative SodaCL YAML files. It runs data contracts, schema validations, and SQL metrics across Snowflake, BigQuery, PostgreSQL, and Spark without requiring transformation code rewrites.
- Spark testing with PyDeequ: Built on Amazon Deequ, PyDeequ executes distributed data quality verifications directly on PySpark DataFrames, calculating metrics using Spark primitives on large-scale datasets.
- Databricks Delta Live Tables: DLT allows developers to define expectations directly inside pipeline code using expect_or_drop or expect_or_fail decorators to filter or halt invalid streaming records.
- Cloud-native services: Google Cloud Dataplex evaluates declarative quality rules and schedules automatic profiling across BigQuery tables and Cloud Storage files.
Warehouse transformations -> dbt tests, dbt-expectations, Elementary, Soda Core Large-scale Spark jobs -> PyDeequ, Amazon Deequ Managed lakehouse (DLT) -> Delta Live Tables expect rules Managed cloud catalogs -> Google Cloud Dataplex, AWS Glue Data Quality
Selecting an assertion tool requires evaluating execution locality:
- Run checks as close to where transformations execute as possible to avoid unnecessary network transfer.
- Warehouse-native checks keep processing inside the database engine, reducing execution overhead and compute costs.