Data profiling is exploratory analysis of a dataset's shape and statistics so you understand what is actually there before you model or trust it.
You are not asserting pass/fail yet. You are discovering distributions, null rates, cardinalities, and surprises.
Typical profile metrics
- Row count
- Null % per column
- Distinct count / cardinality
- Min / max / mean for numerics
- Top-N frequent values for categoricals
- Example patterns (date formats, ID shapes)
orders profile (sample): rows: 1,204,332 order_id null%: 0 email null%: 12% <-- unexpected for "required" field status top: placed 70%, shipped 25%, weird "SHIPPED " 1% amount min/max: -5.00 / 999999 <-- negatives + outliers
When you use it
- Onboarding a new source
- Writing the first dbt/GX tests (tests need realistic thresholds)
- Debugging a sudden metric swing
- Designing grain and SCD rules
Profiling vs testing
Profiling discovers; testing enforces. Good tests often start as profile insights turned into thresholds.
Interview tip: "Profiling = understand the data; tests = enforce rules." Mention null rates, cardinality, and value distributions.