Tutorials › Generative AI Architecture › Fine-Tuning vs. RAG vs. Prompting

Generative AI Architecture · Part 6 of 9

Fine-Tuning vs. RAG vs. Prompting

Three ways to change how a model behaves, solving three different problems.

Three separate tools exist for shaping what a model does: writing a better prompt, retrieving relevant information at query time, and adapting the model's own weights through fine-tuning. They get discussed as competing options, but they answer different questions, and most production systems end up using more than one at once.

The question each one answers

ApproachUse it when…What it can't fix
PromptingThe model already knows how to do the task; it just needs clearer instructions, better structure, or a few worked examples to do it consistentlyThe model has no knowledge of information it was never trained on
Retrieval-augmented generationThe model needs access to information that's private, or that changes too often to have been in its training dataA model that lacks a skill outright, or consistently ignores instructions, no matter how much relevant context it's handed
Fine-tuningThe task needs a specialized, highly consistent behavior baked in: a specific output format, tone, or domain skill that has to hold reliably across thousands of callsFacts that change after the fine-tuning data was collected; a fine-tuned model's knowledge is just as frozen as a base model's

The three also differ in what they cost to adopt. Prompting costs nothing extra to deploy: no training run, no data pipeline, only changed text. RAG adds a data pipeline (the ingestion, chunking, and indexing work) and still no training run. Fine-tuning adds a training run on top of both, computing gradients over a curated dataset, which makes it the slowest of the three to iterate on.

A useful ordering. Reach for prompting first, add RAG when the missing piece is information rather than behavior, and consider fine-tuning only once prompting and RAG together still don't produce behavior consistent enough for the task. Skipping straight to fine-tuning to solve a problem that better instructions or a retrieved document would have solved is usually paying for a training pipeline you didn't need yet.

How they combine in practice

A common production pattern uses all three together: a model fine-tuned to reliably produce a specific output format and tone for a narrow domain, given RAG context at query time for any fact that changes daily, driven by a carefully engineered prompt that tells it how to use both. A fine-tuned customer support model might consistently produce on-brand, correctly formatted responses (the fine-tuning), while being handed the customer's order status from a database each time it answers (the retrieval), inside a prompt that specifies exactly how to phrase a refund policy (the prompting).