Tutorials › Generative AI Architecture › Retrieval-Augmented Generation

Generative AI Architecture · Part 4 of 9

Retrieval-Augmented Generation

Giving a model facts it was never trained on, at the moment it needs them.

A foundation model's knowledge is frozen at training time and general by design; it has no way to know about a company's internal documentation, yesterday's support tickets, or a database row that changed an hour ago. Retrieval-augmented generation, RAG, closes that gap by finding the small number of relevant documents for a given question and handing them to the model as part of the prompt, so the model answers from what it's just been given instead of from what it happened to memorize during training. Semantic search supplies the retrieval half; a production pipeline wraps several more steps around it.

The full pipeline

The pipeline splits into two phases running on different schedules.

Ingestion pipeline: runs ahead of time

flowchart TD
  A[Ingestion] --> B[Parsing]
  B --> C[Chunking]
  C --> D[Embedding]
  D --> E[Indexing]
  

Query-time pipeline: runs on every query

flowchart TD
  Q[Query] --> RET[Retrieval]
  IDX[Indexing] --> RET
  RET --> RANK[Reranking]
  RANK --> PROMPT[Prompt construction]
  PROMPT --> GEN[Generation]
  GEN --> VAL[Validation]
  

Chunking: size, overlap, and metadata

Chunk size is the length, in tokens or characters, of each piece a document is split into before embedding. Too large, and a chunk mixes several unrelated topics into one vector, diluting what it's about and making it a worse match for a narrow question. Too small, and a chunk loses the surrounding context that gave it meaning: a sentence that only makes sense next to the paragraph before it.

Chunk overlap is deliberately repeating a small amount of text at the boundary between adjacent chunks, so a fact split across a chunk boundary still appears intact in at least one of them. Metadata attached to each chunk (source document, section title, date, author, access permissions) doesn't change what gets embedded, but it's what makes filtering possible later: retrieving only chunks from documents a specific user is allowed to see, or only ones updated in the last month.

Hybrid search and reranking

Because keyword search and semantic search fail in different ways, many production systems run both and combine the results, an approach called hybrid search: keyword search catches exact terms semantic search can miss (product codes, error messages, names), and semantic search catches paraphrases keyword search can't see. Reranking then reorders that combined candidate list, often dozens of chunks, since the retrieval pass that produced it is tuned for speed over precision.

Context pollution. Every retrieved chunk placed in the prompt takes up part of a limited context window and can distract the model even when it's not directly wrong. A handful of highly relevant chunks reliably outperforms a large pile of loosely relevant ones.

Where RAG fails

A few failure modes show up often enough to plan for from the start:

What doesn't belong in a vector database

It's tempting, once a RAG pipeline exists, to route every kind of data lookup through it. Structured, relational data, user accounts, orders, inventory counts, anything with exact fields and precise relationships, belongs in a relational database and is queried with SQL. “What's this customer's current subscription tier?” has one correct answer sitting in a row; a vector search would return the closest-sounding match. Databases: Relational and NoSQL covers how to choose the right store for structured data. RAG handles unstructured, meaning-driven lookups (a policy document, a past conversation), and most applications run both side by side.