What RAG is and why it is the cheapest way for AI to know your business
You do not need to train a model on your data. In the overwhelming majority of cases, handing over the right document alongside the question is enough.
RAG stands for retrieval augmented generation. Instead of hoping the model knows something about your company, the system first searches your documents for the relevant passages and then hands those passages over with the question, asking for an answer based only on them. It is cheaper than training a model, it updates the moment you swap a document, and it lets you show where each answer came from.
What you get from this article
- RAG searches first and answers after: the model reads your material instead of trying to remember it.
- Updating the base means swapping a file, not retraining anything.
- The answer can cite the source passage, which makes checking fast.
- RAG quality depends far more on document preparation than on the model you chose.
- It does not answer a question that requires counting, summing or cross-referencing the whole base.
- Without permission control in the search, RAG becomes an internal document leak.
The problem RAG solves
A language model was trained on public text up to a certain date. It does not know your price list, your standard contract, your equipment manual or your customer history. Asking a bare model about that produces an invented answer with the appearance of correctness, because it fills the gap with what would be plausible for a company in your sector.
The intuitive reaction on discovering that is to want to train the model on your own data. That is the most expensive, slowest and, in the overwhelming majority of cases, the wrong solution. It requires preparing thousands of examples, costs a lot, has to be redone when the information changes, and still does not guarantee the model recalls a specific detail correctly.
RAG solves it by another route, and the route is almost obvious once you hear it: instead of making the model know, make it read. With every question, the system searches your base for the passages related to it, pastes those passages into the request and asks for an answer based on them.
The closest analogy is a new person on the team with access to the filing cabinet. They have not memorized the manual, and they do not need to: when a question arrives, they find the right page and answer from there. It is more reliable than memorizing, and far easier to keep up to date.
Instead of making the model know, make it read. Almost every project that starts out wanting to train an AI actually needed this.
How it works, step by step
The system has two halves. One happens beforehand, once per document; the other happens with every question.
Beforehand, the preparation. Your documents are broken into manageable chunks, normally one to three paragraphs. Each chunk is converted into a list of numbers representing its meaning, called a vector. Those vectors go into a database specialized in finding the vector most similar to another. Passages about similar things end up with similar vectors, even without sharing the same words.
At question time. The question also becomes a vector. The database returns the closest passages, typically between three and ten. Those passages are pasted into the request along with an instruction such as "answer using only the material below and cite where you got it". The model answers. Done.
The practical advantage this architecture gives is enormous and rarely advertised: because you control what goes into the request, you control the answer. Swap the document and the answer changes on the next question. Remove a file and it stops being used immediately. There is no retraining, no wait, no tuning cost.
What this solves in practice
Picture a building materials distributor with a catalog of 4,000 items, a price list that changes every week and a sales team that spends the day answering the same questions on WhatsApp. An assistant with RAG over the catalog and the commercial rules answers availability, product equivalence and payment terms without anybody memorizing anything, and the update is the same file that was already generated every week.
Or an accounting firm with hundreds of pages of internal procedure. A junior assistant question about which schedule applies to a specific client is exactly the kind of question where the material exists, is written down, and nobody can find it.
The common pattern across the cases that work is this: the answer already exists in writing somewhere, and the problem is location, not reasoning. When the problem has that shape, RAG is the right tool and the return shows up fast.
And there is a side benefit that tends to be the most valued after a few months: the questions the system cannot answer become a report of what is missing from your documentation. That reveals gaps nobody was mapping.
What RAG does not solve
It is worth being specific here, because the wrong expectation is what kills projects.
A question that requires the whole base. "How many contracts expire in March?" is not a RAG question. The search brings the passages most similar to the question, not every contract. Counting, summing and cross-referencing are the work of a database and a structured query. Many people try to force that into RAG and receive an invented number that looks right.
Information that is not written down. If the rule lives in the manager head and was never documented, RAG has nothing to find. It does not deduce policy from history. A considerable part of the real work in a RAG project is writing down what was never written.
Documents the machine cannot read. A PDF that is a photo of a scanned page, a spreadsheet with information encoded in cell colour, a contract with the important clause inside an image. All of that has to go through text recognition first, and the quality of that stage sets the ceiling for everything after it.
And it reduces, but does not eliminate, hallucination. The model can still misread a passage or join two pieces that should not be joined. Requiring the passage to be cited is what makes that error easy to catch.
Where the build usually goes wrong
A RAG that performs badly is almost never a model problem. It is a preparation problem, and always at the same points.
Chunk size. Too large brings a lot of noise along with the answer and confuses. Too small cuts the information in half, and the answer arrives without the condition that was in the next paragraph. The adjustment is empirical, done with real questions, and it is the parameter with the biggest impact on the final result.
Chunks losing their context. A passage saying "the term is 30 days" is useless on its own: 30 days for what, in which contract, under what condition. Each chunk has to carry where it came from, which document, which section, which version. Without that, the search finds and the answer misses.
Search purely by meaning. Vectors are great at finding a similar subject and bad at finding an exact code. If somebody asks about part "XR-4410", a meaning-based search can return similar parts instead of that one. The known solution is combining meaning search with literal word search, and reranking the results of both.
A dirty base. Three versions of the same contract, two of them old. The system does not know which one counts and cites the wrong one with the same confidence. Before anything technical, somebody has to decide which document is the valid version. That is the most tedious work in the project and the most decisive.
- Chunks of one to three paragraphs, tuned with real questions rather than guesswork.
- Each chunk carries its document, section, version and date of origin.
- Combined search: by meaning and by literal word, with the results reranked.
- One valid version per document, decided by a person before indexing.
- The answer required to cite the passage, for fast checking.
The security part almost every project forgets
This is the most serious risk of a badly built RAG, and it shows up late, once the system is already in use.
If the search does not respect permissions, anyone talking to the assistant can extract anything in the base. An employee asks about the holiday policy and, if the base contains the salary spreadsheet, the passage can come along. That is not a model failure: it is an architecture failure. Permissions have to be applied in the search, filtering what can be retrieved for that user, before the model sees anything.
Filtering afterwards does not work. Instructing the model not to talk about salaries is a fragile barrier, worked around by a well-phrased question, and it breaks silently. If the passage reached the model, it already left your control.
There is also a less obvious risk: if your base indexes documents that came from outside, such as customer email or web pages, malicious text planted there can contain instructions addressed to the model. Retrieved content is data, never a command, and the system has to be built assuming that.
What it costs and where to start
RAG is the cheap alternative, and it is worth understanding why. There is no training, so there is none of that cost. What exists is the cost of generating the vectors once per document, which is low, and the cost of each question, which is the text of the retrieved passages plus the question and the answer.
The arithmetic changes once you realize the retrieved passages go into the request with every question. Bringing ten large passages instead of four precise ones multiplies the cost of every question and makes the answer worse too. Precise retrieval is saving and quality at the same time, which is rare. The details of how that cost accumulates are in tokens and the context window.
Where to start, in the order that avoids waste: pick one set of questions the team answers often and whose answers already exist in writing. Gather only the documents that answer those questions. Write twenty real questions with the correct answer beside each, produced by somebody who knows the subject. Build the simplest possible system and measure it against those twenty.
Those twenty questions are the most valuable item in the whole project, and they are what almost everybody skips. Without them, nobody can say whether a change improved or worsened anything, and the project runs on opinion. If you want to build this in your company with that rigour from the start, that is how ROO3 AI consulting structures the first delivery.
Frequently asked questions
What does RAG stand for?
RAG stands for retrieval augmented generation. The name describes the order of operations: the system first retrieves relevant passages from your documents and only then generates the answer, using those passages as its basis.
What is the difference between RAG and training a model on my data?
Training alters the model and is expensive, slow and has to be redone when the information changes. RAG leaves the model untouched and hands the document over with the question. Updating means swapping a file, and the answer can cite its source, which training does not offer.
Do my documents go inside the model?
No. They stay in your base and only the relevant passages are sent with each question. Whether the passage sent will be used in the provider future training depends on the service contract, and business plans normally exclude that. It is worth confirming before indexing sensitive data.
Is RAG good for asking how many customers I have?
No. Counting, summing and cross-referencing require a structured query against a database, not a similarity search. RAG brings the passages most similar to the question, not the whole base, and forcing that kind of question produces invented numbers that look correct.
Do I need an expensive model to do RAG?
Usually not. Because the material comes with the question, the task is reading and summarizing rather than knowing, and mid-tier models handle it well. The money goes further invested in document preparation and search quality than in swapping models.
How long does it take to build a RAG?
A first useful version over a small, well-defined set of documents takes weeks, not months. What stretches the timeline is almost always cleaning the base: deciding which document is the valid version, converting whatever is an image, and writing down what was never documented.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

RAG, fine-tuning or prompt: which to use in each case
When to use a prompt, when to use RAG and when fine-tuning is genuinely justified. The question that decides, the cost...
10 min read
AI hallucination: why it invents and how to reduce it
What AI hallucination is, why it happens by design, what the research shows about the cause, and the techniques that...
10 min read
Tokens and the context window: why the AI bill comes in high
What a token is, what a context window is and how billing really works. With the arithmetic done and the four mistakes...
11 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.