AI hallucination: why the model invents so confidently, and what actually reduces it
It is not a bug, it is not missing data, and it is not solved by asking the AI not to make things up. The cause lies in how these systems are trained and evaluated.
Hallucination is when the model produces false information with the same fluency and the same confidence it uses for correct information. The cause is not missing data: OpenAI researchers showed in 2025 that standard training and evaluation methods reward guessing and punish "I do not know", so the model learns to always take a chance. Reducing it depends on architecture, handing over the source alongside the question and verifying the output, not on politely asking it not to invent.
What you get from this article
- Hallucination comes from the design of the system, not from a defect a better model solves on its own.
- Benchmarks that score zero for "I do not know" teach the model to guess, according to OpenAI research.
- The highest-impact technique is handing over the source with the question, instead of trusting the model memory.
- Asking "do not invent" in the prompt has a small effect and gives a false sense of safety.
- Verification by code is worth more than verification by another model when the data is checkable.
- Proper nouns, dates, numbers and citations are the four categories that fail most.
What it is, and why the name gets in the way
Hallucination is the name that stuck for the moment when the model states something false as naturally as it states something true. A statute that does not exist, a wrong date, an invented scientific study with a plausible author and year, a feature that product never had.
The name gets in the way because it suggests an altered state, an exception, something outside the normal. It is not. From the system point of view, nothing different happened. It performed exactly the same operation as always, predicting the next chunk of text, with the same confidence as always. What separates a hit from a miss is you, from the outside.
That is why there is no warning. There is no internal meter that fires "careful, I made this part up". The voice is the same, the sentence structure is the same, the fluency is the same. A more precise term would be confabulation: filling a gap with something plausible without noticing that you are filling it.
Understanding that part changes how you use the tool. The error comes with no signal, so verification cannot depend on you noticing something odd. It has to be systematic at the points where being wrong is expensive.
The deeper cause: training rewards guessing
In September 2025, OpenAI researchers published work that gives the most direct and most uncomfortable explanation of the phenomenon. The title is Why Language Models Hallucinate, and the central argument is about incentives, not capability.
The logic is that of a multiple-choice exam with no penalty for wrong answers. If leaving it blank scores zero and guessing carries some chance of being right, the rational student guesses at everything. The benchmarks used to evaluate and compare models almost all work that way: a correct answer scores, a wrong answer does not score, and "I do not know" does not score either. From the point of view of the score, guessing dominates admitting ignorance.
Because those benchmarks guide the development and public comparison of models, the result is a system trained to always take a chance. Excessive confidence is not an accident: it is rewarded behaviour. The authors propose changing evaluation criteria to give partial credit for well-placed uncertainty and to penalize confident error.
The practical conclusion that matters to anyone buying: hallucination is not a bug the next version solves. New models hallucinate less because the base improved, and they carry on hallucinating because the incentive remains. Anyone promising AI that never gets things wrong is selling something that does not exist.
In an exam with no penalty for wrong answers, guessing always beats leaving it blank. That is exactly the incentive we train into these models.
Why it is always the same four things
There is a clear pattern in what fails, and knowing it lets you aim your checking rather than rereading everything.
Proper nouns and references is the leading category. Book authors, study names, case numbers, clauses in technical standards. The model learned the shape of a citation, and the shape is easy to reproduce without the correct content. The result is a reference that looks impeccable and does not exist.
Dates and numbers come next. The year of a law, a percentage, an amount, a product version. These vary a lot between similar texts, so the statistical pattern is weak exactly where precision matters.
Specific details about a product or company is the third. If the information was not sufficiently present in the training material, the model fills in with what would be reasonable for a company of that type. And anything recent is the fourth: everything after the training cut-off is reconstructed by analogy, unless the system searches live.
Notice what those four categories have in common: they are exactly what gets copied out of an answer into a document without rereading. That is why the damage is disproportionate to the number of errors.
What genuinely reduces it, in order of effect
The highest-impact technique, and it is not close, is handing over the information with the question. Instead of asking the model what your company returns policy says, you send the policy in the request itself and ask for an answer based on it. The job stops being remembering and becomes reading, and reading is what it does well.
Doing that automatically and at scale has a name: RAG. A system searches your base for the relevant passages, pastes them into the request, and the model answers using only that. It does not eliminate the error, because the model can still misread or extrapolate, but it changes the nature of the problem: from free invention to reading material you control.
The second technique is requiring the passage to be cited. Ask for every claim to come with the exact piece of the document it rests on. That helps in two ways: it makes free invention harder and, above all, it makes checking fast, because the reviewer does not have to search, only compare.
The third is explicitly authorizing "I do not know". Instructions such as "if the information is not in the documents provided, say it is not available and stop" work better than a generic ban on inventing, because they offer an acceptable way out instead of only closing a door.
- High effect: sending the source with the question, instead of trusting the model memory.
- High effect: requiring the literal passage that supports each claim.
- Medium effect: authorizing "it is not in the material provided" as a valid answer.
- Medium effect: validating in code what is checkable, such as tax numbers, dates, amounts and whether a record exists.
- Low effect: asking the prompt not to invent, without changing anything else.
What barely works, and everybody tries
It is worth a paragraph each, because those three tactics dominate the tutorials and give a sense of safety that does not match the result.
Asking it not to invent. Putting "do not invent information" in the prompt has a small effect. The model is not inventing on purpose: it does not know it is. The instruction asks for a behaviour that depends on an awareness the system does not have.
Asking whether it is sure. When you question it, it frequently changes the answer, and that looks like self-correction. It is not. It is the effect of training that rewards agreeing with the user. It changes just as easily when it was right, which is worse: you just destroyed a good answer.
Asking the model to review itself. It helps with coherence and logic errors, and it helps little with facts, because the second pass has exactly the same limitation as the first. Where the data is verifiable by code or by querying a system, verifying by code is worth far more than verifying by another model.
How to decide where this matters in your company
Not every task demands the same rigour, and treating everything with maximum rigour kills the gain. The right question is not whether the model can be wrong, it is how much each error costs and who notices.
There are tasks where hallucination barely matters, because all the necessary information is in the input and the output is checked immediately. Summarizing a meeting whose audio is right there, rewriting text you are going to reread, suggesting titles, tidying a messy list. The error shows up in front of whoever asked.
And there are tasks where hallucination is unacceptable without a guardrail: anything involving amounts, contractual deadlines, health guidance, legal information, or text that goes straight to the customer without passing anybody. Here the model comes in as a draft and the decision stays with a person, or the data comes out of a system and the model only writes around it.
That is why at ROO3 one project rule is fixed: amounts are not calculated by a language model. The arithmetic is done in code, which is deterministic and testable, and the model comes in to explain the result in plain language. Getting money wrong is the category of error that no time saving makes up for.
A system design that survives the error
A reliable system with AI is not the one that prevents errors, it is the one that survives them. The difference sits in three layers working together.
At the input: the model receives the necessary material instead of depending on memory, and the material arrives trimmed to what matters. A whole document in the window is an invitation for the model to lose itself in the middle and fill the gap with invention.
At the output: everything checkable by machine is checked by machine. If the model returned a company registration number, validate the format and the existence. If it returned an amount, recalculate it. If it returned a reference to a document, check whether that passage really exists in the document. That verification costs almost nothing and catches most of the serious errors.
In the flow: one human checkpoint positioned where the cost of error is high, not at every stage. Reviewing everything cancels the gain and tires the person, who starts approving on autopilot after two weeks. Reviewing what matters keeps the attention where it is worth having.
If you are designing a process like that now and want to define where those guardrails sit in your case, that is exactly the conversation of ROO3 AI consulting. And if the problem is choosing the model, the public measurements are gathered in the AI Benchmark.
Frequently asked questions
Why does AI invent instead of saying it does not know?
Because training and evaluation reward guessing. OpenAI researchers showed in 2025 that the benchmarks used to compare models score zero both for a wrong answer and for "I do not know", which makes taking a chance always better than admitting ignorance from the point of view of the score.
Do newer models hallucinate less?
They hallucinate less, and they carry on hallucinating. The knowledge base improves and the reasoning improves, but the incentive that rewards confident answers remains. No version solves the problem completely, so the design of the system around it is still necessary.
Does asking the AI in the prompt not to invent help?
It helps little. The model does not know it is inventing, so the instruction asks for a behaviour that depends on a perception it does not have. It works better to explicitly authorize the answer "it is not in the material provided" and to hand the material over with the question.
Does RAG eliminate hallucination?
It does not eliminate it, it changes its nature. With RAG the model answers from passages you control, which greatly reduces free invention, but it can still misread a passage or extrapolate beyond what is written. It is the highest-impact technique, not a guarantee.
Where do the errors show up most often?
In four categories: proper nouns and references, such as authors and case numbers; dates and numbers, such as the year of a law and percentages; specific details about a product or company; and anything too recent to be in the training. Those are precisely the items that tend to be copied without rereading.
Does asking "are you sure?" correct the answer?
Not reliably. The model tends to change its answer when questioned because training rewards agreeing with the user, and it changes just as easily when the original answer was right. Questioning can destroy a good answer.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

What RAG is and why it is the cheapest way to use AI
What RAG is, explained practically: how the search-before-answering works, when it solves the problem, what it does not...
10 min read
What an LLM is: how it works inside and what it does not do
What an LLM is, explained without maths: how the model predicts the next word, why that works so well, and which limits...
12 min read
How to write a good prompt: the method that works
The method for writing prompts that work: what to include, in what order, and a comparison of bad and good requests for...
10 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.