RAG, fine-tuning or prompt: which to choose, with the cost and effort of each path
Three routes to making AI work with your business. Most projects choose the most expensive one because they do not know the question that separates the three.
The question that decides is what is missing. If instruction is missing, the problem is solved in the prompt. If information that changes over time is missing, the answer is RAG. If what is missing is consistent behaviour you cannot describe in text, such as a very specific format or a style of your own, then fine-tuning is justified. In practice, most business cases stop at the prompt or at RAG, and fine-tuning is the right choice in a small minority.
What you get from this article
- A prompt fixes missing instruction, RAG fixes missing information, fine-tuning fixes missing behaviour.
- Fine-tuning is not the way to teach facts to a model: for facts, RAG wins on cost, timeline and auditability.
- Fine-tuning has to be redone when the base model changes, and the base model changes often.
- The right order is prompt, then RAG, then fine-tuning, measuring at each stage.
- Without a set of questions with correct answers, none of the three options can be evaluated.
- The three combine: fine-tuning for format with RAG for content is a common design.
The question that separates the three
Before comparing technologies, it is worth asking a question that settles most doubts: what exactly is missing from the answer you received?
If the model answered with the right information but in the wrong format, the wrong tone, too long, or without following the step you wanted, instruction is missing. That is solved in the prompt, and it is free.
If the model answered in a well-structured way but with information that is not your company, or invented a figure that only exists in your documents, information is missing. That is solved with RAG, handing the material over with the question.
If the model understands the task, has the information and still does not get the manner right consistently, and you have already tried explaining it in the prompt several ways without success, then behaviour is missing. That is the only case where fine-tuning is the right answer.
Almost every project that starts by asking "how do I train an AI on my data?" is actually in the second case, which is the cheapest of the three.
Is instruction missing, information missing or behaviour missing? The answer to that question is the whole decision.
Prompt: the option that solves more than it seems
A prompt is the instruction you send with the question, and it includes examples. It is the most underrated option because it is free and does not look like technology, and it is where most of the gain sits in most cases.
The technique that pays off most here is giving examples. Two or three inputs with the exact output you want teach format, tone and level of detail better than three paragraphs of description. It is the difference between explaining how to write and showing a finished text.
The practical advantage of the prompt is the speed of iteration. You change it, test it and see the result in seconds, with no training cost and no waiting. Neither of the other two options offers that, and it matters more than it seems when you still do not know exactly what you want.
The limit appears when the instruction becomes enormous. A prompt several pages long is expensive, because it travels in every call, and it becomes fragile, because too many instructions start competing with each other and the model begins ignoring some of them. When you reach that point, it is a sign you are overdue to look at the other options. The techniques are detailed in how to write a good prompt.
RAG: when what is missing is information
RAG searches your documents for the relevant passages and hands them over with the question. It is the right choice whenever the answer depends on information that is yours, that changes, or that is too large to fit in a fixed instruction.
Three characteristics make RAG hard to beat in that scenario. The first is updating: change the document, change the answer on the next question, with no retraining and no waiting. The second is auditability: the answer can cite its source passage, which makes checking fast and lets you discuss the answer with whoever knows the subject.
The third is access control. Because the search happens first, you filter what that user can see before the model receives anything. In a fine-tuned model, the information is diluted into the weights and no filter is possible: whoever has access to the model has access to everything that went into it.
That last point usually decides the choice in companies with sensitive data, and it is the least cited argument in tutorials. How RAG works inside is in what RAG is.
Fine-tuning: what it actually does
Fine-tuning is taking a finished model and continuing to train it with your examples, adjusting the weights. The result is a model that behaves a specific way by default, without you having to explain it every time.
The most expensive misunderstanding in the market is here: fine-tuning is not the way to teach facts to a model. It adjusts behaviour, style and format. Specific facts go in diffusely and unreliably, and you cannot know whether a fact was genuinely learned nor correct it later without retraining. For facts, RAG wins on every dimension that matters.
Where it shines is in three situations. A very rigid, repetitive output format that has to come out identical thousands of times. A tone and style of your own, hard to describe in words but easy to show in a hundred examples. And cost reduction: a small model tuned for a narrow task can reach the quality of a large model on that specific task, for a fraction of the price and with less waiting.
That third case is the one that pays off most in production, and it is precisely the least talked about. If you run the same mechanical task a million times a month, tuning a small model for it can transform the economics of the whole project.
The real cost of each path
Comparing only the training price misleads, because the dominant cost of fine-tuning is not that. It is preparation and maintenance.
With a prompt, the cost is your time writing and testing, plus the instruction tokens in every call. Iteration in minutes. Maintenance almost nil: the rule changed, you edit some text.
With RAG, the cost is building the search infrastructure, preparing the documents and keeping the base clean. The expensive, tedious part is the cleaning: deciding which version of each document counts, converting whatever is an image, and writing down what was never documented. Continuous maintenance, but easy to delegate.
With fine-tuning, the cost starts with preparing hundreds to thousands of quality examples, which is specialist human work and cannot be outsourced cheaply. Then comes the training, the testing and the hosting. And then the maintenance almost nobody budgets for: when the base model gets a new version, your tuning does not come with it. Either you stay stuck on an ageing model, or you redo the process. Base models change often, and that is a recurring cost disguised as a one-off investment.
- Prompt: hours of work, iteration in minutes, maintenance by editing text.
- RAG: weeks of building, with more cost in cleaning the documents than in the technology.
- Fine-tuning: hundreds to thousands of examples prepared by people who know the subject, plus rework with every base model change.
The right order, and why it saves money
The sequence that works is always the same, and it is designed so you discover early that you do not need the next stage.
Step zero, and none of the three works without it: build a set of twenty to fifty real questions with the correct answer beside each, written by somebody who knows the subject. Without that, you have no way of knowing whether a change improved or worsened anything, and the project runs on the opinion of whoever spoke last.
Step one: solve it with a prompt, including examples. Measure against the set. A lot that looked like it needed training stops here, and stops in an afternoon.
Step two: if what is missing is your own information, build the RAG. Measure again. Most business projects end here, with a good result and a predictable cost.
Step three: only if a consistent behaviour problem remains, or there is a clear need to cut cost at high volume, evaluate fine-tuning. And go into it knowing you have taken on recurring maintenance.
All three together, which is how it usually ends
In practice, mature systems do not choose one: they use each piece for what it does best, and the design becomes clear once you separate the three questions.
A concrete example of how that combines: a company issuing thousands of standardized reports a month can use fine-tuning so the report format comes out exactly the same every time, RAG to bring in the technical standards applicable to each case, and a prompt to adjust the level of detail to the recipient. Three layers, three different problems.
The symmetrical mistake also exists and is expensive: using RAG for everything, including things a prompt would solve, adds latency, cost and complexity with no gain. Not every question needs a search.
And it is worth recording the rule that avoids most of the waste: do not start with the most expensive path out of fear that the cheapest will not work. Testing the prompt costs an afternoon. Discovering after three months of fine-tuning that the prompt would have done it costs the project.
If you are at that decision now and want somebody to look at the concrete case before you commit budget, that is exactly what ROO3 AI consulting does in the first stage: defining which of the three questions is yours, before writing a single line of code.
Frequently asked questions
What is the difference between RAG and fine-tuning?
RAG hands information over with the question and leaves the model untouched, so updating means swapping a file and the answer can cite its source. Fine-tuning alters the model weights to change behaviour and style, and it is not a reliable way to teach specific facts.
Does fine-tuning teach facts to a model?
Not reliably. Facts go in diffusely, you cannot verify whether a specific fact was learned nor correct it without retraining, and the answer cites no source. For facts, RAG is better on cost, timeline and auditability.
When is fine-tuning worth it?
In three cases: a very rigid, repetitive output format; a tone or style hard to describe in words but easy to show in examples; and cost reduction at high volume, by tuning a small model for a narrow task instead of using a large model.
How many examples do I need for fine-tuning?
It depends on the task, but the requirement usually starts in the hundreds and reaches thousands of good quality examples. The dominant cost is preparing those examples, which requires people who know the subject, rather than the training itself.
Do I have to redo fine-tuning when a new model ships?
Yes, if you want to use the new model. The tuning is done on a specific base model and does not migrate automatically. Because base models change often, that is a recurring cost that has to go into the budget from the start.
Can you use all three together?
Yes, and mature systems generally do. A common design is fine-tuning handling the output format, RAG bringing in the up-to-date content and a prompt adjusting the tone to the case. Each layer solves a different kind of gap.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

What RAG is and why it is the cheapest way to use AI
What RAG is, explained practically: how the search-before-answering works, when it solves the problem, what it does not...
10 min read
How to write a good prompt: the method that works
The method for writing prompts that work: what to include, in what order, and a comparison of bad and good requests for...
10 min read
Tokens and the context window: why the AI bill comes in high
What a token is, what a context window is and how billing really works. With the arithmetic done and the four mistakes...
11 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.