A vector database stores embeddings, which are lists of numbers that represent the meaning of text, images or other data, and finds the ones closest to a query vector. Items with similar meaning end up near each other, so searching by closeness gives semantic search.
What it does
Take a product description and convert it to a vector of, say, 768 numbers with an embedding model. Store these vectors with an id and some metadata. At search time, embed the user's question the same way, and ask the database for the 10 nearest vectors by cosine similarity or distance. The results are the items whose meaning is closest, even when they share no words with the question.
Why it is "approximate"
Comparing the query with every vector exactly works for thousands of items but is too slow for hundreds of millions. Vector databases use approximate nearest neighbour (ANN) indexes that trade a small loss in recall for big speedups. HNSW builds a layered graph that you walk toward the query. IVF splits the space into clusters and searches only the nearest few. Settings control the trade-off between speed, memory and recall.
Metadata filtering
Real queries combine meaning with conditions: "similar documents, but only from the finance team and from 2024". Good systems support filters together with vector search. Plan which metadata fields you need before building.
Options
- Add-ons to databases you already run: pgvector for Postgres, and vector search in BigQuery, Snowflake and Databricks. Convenient, with one system and shared governance. Fine for moderate scale.
- Dedicated systems: Pinecone, Weaviate, Milvus, Qdrant. Built for large scale and low latency, with more features for vectors, and another system to run and secure.
When a data engineer needs one
For semantic search, recommendations, duplicate detection, and as the retrieval store behind RAG applications. If the data already sits in a warehouse and the scale is modest, try the built-in vector features first.
What it is not
It is not a replacement for the warehouse or a relational database. It does not do joins and aggregations well, and it is rarely the system of record. Your source data lives elsewhere, and the vector store is an index built from it, rebuilt when needed.
Operational points
Embedding model changes need a re-index, indexes use a lot of memory, and results are approximate, so evaluation with test queries is essential.