A data catalog is a searchable inventory of datasets with metadata: what a table means, who owns it, where it lives, how fresh it is, and how to use it safely.
Without a catalog, people Slack "where's the orders table?" and copy the wrong one.
What a good catalog entry has
- Name, description, domain
- Owner / steward
- Schema and column descriptions
- Grain and primary key
- Freshness / last updated
- Lineage links
- Quality status or known caveats
- Access / PII tags
Catalog card: analytics.fct_orders owner: payments-de grain: one row per order_id freshness SLA: < 1 hour quality: 12/12 tests green PII: email hashed; raw email in restricted schema
Catalog vs lineage vs quality
- Catalog: discovery + documentation + ownership
- Lineage: how assets connect
- Quality: whether assets currently meet rules
Modern catalogs often show all three together.
Interview tip: "Catalog helps people find the right trusted table." Mention ownership, descriptions, freshness, and PII tags.