Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Pipelines for LLM and RAG applications

Pipelines & scenarios · Core Concepts Not Yet Covered

Pipelines for LLM and RAG applications

Mediumpipelines-71
ragllmembeddingsvector-searchchunking

Question

What does a data pipeline for a RAG (retrieval-augmented generation) application look like?

Solution

A RAG (retrieval-augmented generation) application answers questions by finding relevant pieces of your own documents and giving them to a language model together with the question. The data pipeline's job is to turn documents into searchable chunks and keep that index correct as documents change.

documents -> parse/clean -> chunk -> embed -> vector index (+ metadata)
question -> embed -> nearest chunks -> prompt to LLM -> answer

The ingestion steps

  • Collect documents from their sources (wikis, PDFs, tickets, drives, databases), with an identifier and a last-modified time for each.
  • Parse and clean: extract text from PDFs and HTML, remove headers, footers, navigation and boilerplate, and keep structure such as titles and tables where it helps. Poor parsing is a leading cause of poor answers.
  • Chunk: split text into pieces, often a few hundred tokens, usually along natural boundaries (sections, paragraphs), with a small overlap so that a sentence cut at the edge is still found. Chunks that are too big dilute meaning, and chunks that are too small lose context.
  • Embed: send each chunk to an embedding model, which returns a vector that captures its meaning.
  • Store: write the vector, the chunk text and metadata (source, document id, section, date, access group, version) to a vector database or a vector index in your warehouse.

Keeping it fresh

Do not re-embed everything on each run. Detect which documents changed (by modified time or content hash), and re-embed only those, replacing their old chunks. Handle deletes: when a document is removed, remove its chunks, otherwise the system will quote things that no longer exist. Handle a changed embedding model carefully, because vectors from different models are not comparable, which means a full rebuild.

Access control

If documents have permissions, store the allowed groups in the metadata and filter at query time, so people only retrieve what they may see. Leaking restricted text through an answer is a serious failure.

Quality and cost

Evaluate retrieval with a set of real questions and the passages that should be found, and measure how often they are returned. Track the cost of embedding calls, since a large corpus re-embedded often adds up. Log which chunks were used for each answer, so you can debug wrong ones.

Where it differs from classic pipelines

The data is unstructured, "correct" is judged by relevance, and new failure modes appear (bad parsing, stale chunks, leaked permissions). Say that the ordinary pipeline skills (idempotency, incremental loads, lineage, access control) still apply fully.

PreviousNext