Accepted answer
dbt docs, the manifest JSON, and exposures for dashboards get you most of the way there. Add OpenLineage and Marquez if you need runtime lineage too.
Startup, no Collibra or Alation budget. Need column-level lineage from raw to dashboard for a compliance ask.
Are dbt docs and manifests enough here? What's a lightweight stack that actually works?
Accepted answer
dbt docs, the manifest JSON, and exposures for dashboards get you most of the way there. Add OpenLineage and Marquez if you need runtime lineage too.
Could you share a sketch of the salting logic?
Careful with NULL in join keys, they'll drop rows in an inner join.
ELI5: think of it like a phone book. If it's sorted by last name and you search by last name, that's fast. Search by first name instead and you're flipping through every page.
Our dimension is slowly changing, does that change the join order?
Another path: push the compute to the warehouse if the data's already there.
Another path: push the compute to the warehouse if the data's already there.
Consider DuckDB or Polars for this size before spinning up a cluster.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.