11
Accepted answer
For the larger or more dynamic ones, move them to a proper source table maintained by whatever system owns that reference data, and reference it as a dbt source instead of a seed. Keep seed for the genuinely small and static.
Using dbt seed for a handful of small lookup CSVs, country codes, status mappings, worked great early on. Now we have 40+ seed files and some are getting uncomfortably large, tens of thousands of rows. Time to stop?
Accepted answer
For the larger or more dynamic ones, move them to a proper source table maintained by whatever system owns that reference data, and reference it as a dbt source instead of a seed. Keep seed for the genuinely small and static.
Small, rarely-changing lookups are exactly what seed is for. The smell is size and change frequency, not the pattern itself. Tens of thousands of rows in a CSV checked into git is where it starts to hurt, diff noise, repo bloat, slow CI.
Sign in to reply.