Generative AI Architecture · Part 4 of 9
Retrieval-Augmented Generation
Giving a model facts it was never trained on, at the moment it needs them.
A foundation model's knowledge is frozen at training time and general by design; it has no way to know about a company's internal documentation, yesterday's support tickets, or a database row that changed an hour ago. Retrieval-augmented generation, RAG, closes that gap by finding the small number of relevant documents for a given question and handing them to the model as part of the prompt, so the model answers from what it's just been given instead of from what it happened to memorize during training. Semantic search supplies the retrieval half; a production pipeline wraps several more steps around it.
The full pipeline
The pipeline splits into two phases running on different schedules.
Ingestion pipeline: runs ahead of time
flowchart TD A[Ingestion] --> B[Parsing] B --> C[Chunking] C --> D[Embedding] D --> E[Indexing]
- Ingestion pulls source content in: PDFs, wiki pages, support tickets, database exports.
- Parsing converts that raw content into clean text, stripping layout noise, extracting tables, handling scanned pages.
- Chunking splits long documents into smaller pieces, since a whole 40-page manual is rarely the right unit to retrieve.
- Embedding converts each chunk into a vector.
- Indexing stores those vectors in a structure built for fast nearest-neighbor search.
Query-time pipeline: runs on every query
flowchart TD Q[Query] --> RET[Retrieval] IDX[Indexing] --> RET RET --> RANK[Reranking] RANK --> PROMPT[Prompt construction] PROMPT --> GEN[Generation] GEN --> VAL[Validation]
- Retrieval takes an incoming query, embeds it the same way, and finds the closest-matching chunks.
- Reranking re-scores that first cheap pass of candidates with a slower, more accurate model, pushing the best matches to the top.
- Prompt construction assembles the final prompt from the system instructions, the user's question, and the retrieved chunks, using the same prompt template discipline as any other call.
- Generation is the model producing an answer from that assembled prompt.
- Validation checks the answer before it reaches the user: does it cite the retrieved sources, does it avoid claiming something the sources don't support, does it match the expected format.
Chunking: size, overlap, and metadata
Chunk size is the length, in tokens or characters, of each piece a document is split into before embedding. Too large, and a chunk mixes several unrelated topics into one vector, diluting what it's about and making it a worse match for a narrow question. Too small, and a chunk loses the surrounding context that gave it meaning: a sentence that only makes sense next to the paragraph before it.
Chunk overlap is deliberately repeating a small amount of text at the boundary between adjacent chunks, so a fact split across a chunk boundary still appears intact in at least one of them. Metadata attached to each chunk (source document, section title, date, author, access permissions) doesn't change what gets embedded, but it's what makes filtering possible later: retrieving only chunks from documents a specific user is allowed to see, or only ones updated in the last month.
Hybrid search and reranking
Because keyword search and semantic search fail in different ways, many production systems run both and combine the results, an approach called hybrid search: keyword search catches exact terms semantic search can miss (product codes, error messages, names), and semantic search catches paraphrases keyword search can't see. Reranking then reorders that combined candidate list, often dozens of chunks, since the retrieval pass that produced it is tuned for speed over precision.
Where RAG fails
A few failure modes show up often enough to plan for from the start:
- Plausible-but-wrong retrieval. A chunk can be topically similar to the query and still be the wrong answer: an old pricing page instead of the current one, a similar-sounding but different product's documentation. Semantic similarity measures relatedness; correctness is a separate question.
- Chunks losing surrounding context. A chunk that reads as a self-contained fact in isolation can depend on a caveat, exception, or condition stated a paragraph earlier that didn't make it into the same chunk.
- Stale indexes. If the ingestion pipeline doesn't run again when source content changes, retrieval keeps confidently returning content that used to be true.
What doesn't belong in a vector database
It's tempting, once a RAG pipeline exists, to route every kind of data lookup through it. Structured, relational data, user accounts, orders, inventory counts, anything with exact fields and precise relationships, belongs in a relational database and is queried with SQL. “What's this customer's current subscription tier?” has one correct answer sitting in a row; a vector search would return the closest-sounding match. Databases: Relational and NoSQL covers how to choose the right store for structured data. RAG handles unstructured, meaning-driven lookups (a policy document, a past conversation), and most applications run both side by side.