Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Vector databases

Pipelines & scenarios · Core Concepts Not Yet Covered

Vector databases

Easypipelines-72
vector-databaseembeddingsannhnswsemantic-search

Question

What is a vector database, and when does a data engineer need one?

Solution

A vector database stores embeddings, which are lists of numbers that represent the meaning of text, images or other data, and finds the ones closest to a query vector. Items with similar meaning end up near each other, so searching by closeness gives semantic search.

What it does

Take a product description and convert it to a vector of, say, 768 numbers with an embedding model. Store these vectors with an id and some metadata. At search time, embed the user's question the same way, and ask the database for the 10 nearest vectors by cosine similarity or distance. The results are the items whose meaning is closest, even when they share no words with the question.

Why it is "approximate"

Comparing the query with every vector exactly works for thousands of items but is too slow for hundreds of millions. Vector databases use approximate nearest neighbour (ANN) indexes that trade a small loss in recall for big speedups. HNSW builds a layered graph that you walk toward the query. IVF splits the space into clusters and searches only the nearest few. Settings control the trade-off between speed, memory and recall.

Metadata filtering

Real queries combine meaning with conditions: "similar documents, but only from the finance team and from 2024". Good systems support filters together with vector search. Plan which metadata fields you need before building.

Options

  • Add-ons to databases you already run: pgvector for Postgres, and vector search in BigQuery, Snowflake and Databricks. Convenient, with one system and shared governance. Fine for moderate scale.
  • Dedicated systems: Pinecone, Weaviate, Milvus, Qdrant. Built for large scale and low latency, with more features for vectors, and another system to run and secure.

When a data engineer needs one

For semantic search, recommendations, duplicate detection, and as the retrieval store behind RAG applications. If the data already sits in a warehouse and the scale is modest, try the built-in vector features first.

What it is not

It is not a replacement for the warehouse or a relational database. It does not do joins and aggregations well, and it is rarely the system of record. Your source data lives elsewhere, and the vector store is an index built from it, rebuilt when needed.

Operational points

Embedding model changes need a re-index, indexes use a lot of memory, and results are approximate, so evaluation with test queries is essential.

PreviousNext