Generative AI Architecture · Part 6 of 9
Fine-Tuning vs. RAG vs. Prompting
Three ways to change how a model behaves, solving three different problems.
Three separate tools exist for shaping what a model does: writing a better prompt, retrieving relevant information at query time, and adapting the model's own weights through fine-tuning. They get discussed as competing options, but they answer different questions, and most production systems end up using more than one at once.
The question each one answers
| Approach | Use it when… | What it can't fix |
|---|---|---|
| Prompting | The model already knows how to do the task; it just needs clearer instructions, better structure, or a few worked examples to do it consistently | The model has no knowledge of information it was never trained on |
| Retrieval-augmented generation | The model needs access to information that's private, or that changes too often to have been in its training data | A model that lacks a skill outright, or consistently ignores instructions, no matter how much relevant context it's handed |
| Fine-tuning | The task needs a specialized, highly consistent behavior baked in: a specific output format, tone, or domain skill that has to hold reliably across thousands of calls | Facts that change after the fine-tuning data was collected; a fine-tuned model's knowledge is just as frozen as a base model's |
The three also differ in what they cost to adopt. Prompting costs nothing extra to deploy: no training run, no data pipeline, only changed text. RAG adds a data pipeline (the ingestion, chunking, and indexing work) and still no training run. Fine-tuning adds a training run on top of both, computing gradients over a curated dataset, which makes it the slowest of the three to iterate on.
How they combine in practice
A common production pattern uses all three together: a model fine-tuned to reliably produce a specific output format and tone for a narrow domain, given RAG context at query time for any fact that changes daily, driven by a carefully engineered prompt that tells it how to use both. A fine-tuned customer support model might consistently produce on-brand, correctly formatted responses (the fine-tuning), while being handed the customer's order status from a database each time it answers (the retrieval), inside a prompt that specifies exactly how to phrase a refund policy (the prompting).