Fundamentals

What an LLM is, how it works inside, and what it definitely does not do

Understanding the mechanism changes how you use the tool. And it explains, once and for all, why it gets dates wrong and writing right.

Rodrigo Fávaro, fundador da ROO3
Rodrigo Fávaro Founder of ROO3
·12 min read
Abstract illustration of a sequence of text blocks linking together, with the next block highlighted in neon green on a black background.
Short answer

LLM stands for large language model. It is a system trained on an enormous amount of text to do one thing: predict which chunk of a word comes next, over and over, until the answer is formed. It does not consult a database of facts and has no notion of what is true. Everything else, writing, summarizing, translating, coding, is a consequence of doing that prediction very well.

What you get from this article

  • An LLM predicts the next chunk of text, and all of its abilities come out of that single operation.
  • It has no database of facts: the knowledge is diluted across billions of numeric parameters.
  • Getting a date or a number wrong is not a bug, it is a direct consequence of the prediction mechanism.
  • Every answer starts from scratch: the model does not remember the earlier conversation unless it is resent.
  • Where the cost of being wrong is high, the model text is a draft, not a decision.
  • The way to improve the result is to give context, not to ask more politely.

The operation that explains everything else

An LLM does one thing, repeated thousands of times per second: given the text that exists so far, it calculates the most likely next chunk, picks one, and starts the calculation again now counting the chunk it just wrote. That is it. There is no hidden second function.

Almost everyone first reacts by thinking that is far too little to explain what the tool does. That is a reasonable reaction, and it is wrong for a specific reason: to predict the next word well across any text in the world, the system is forced to learn an enormous amount of structure. To complete "the capital of France is", it has to have captured geography. To correctly complete the end of a contract, it has to have captured legal form. To close a code function, it has to have captured logic.

That answers a question that always comes up: does the model understand what it is saying? It depends what you call understanding, and that discussion is genuinely philosophical, not rhetorical. From the practical point of view, which is what matters to anyone about to use it, the answer is different: it behaves like something that understands across an enormous range of tasks, and it fails in a way nothing that understood would fail.

Hold on to that last sentence. An LLM error pattern is different from a human error pattern, and that is why it surprises people. It writes a well-structured legal opinion and gets the date of a statute wrong. A person who gets that date wrong would normally also write a bad opinion. Here, the two things are independent.

None of its abilities was programmed by anybody. All of them are a side effect of doing a single prediction very well.

Where the knowledge is kept

There is no database inside the model that you can open and query. There is no table with the capital of every country and no file with the text of statutes. There is a gigantic quantity of numbers, called parameters, adjusted during training.

Those parameters encode statistical relationships between chunks of text. The knowledge is spread across them, not stored at an address. That is why nobody can delete a specific fact from a finished model, and why the model cannot say where it got a piece of information: it did not take it from a place, it reconstructed it from a distributed pattern.

Here is the root of the behaviour that most annoys new users. When the pattern is strong, because that thing appeared thousands of times in similar texts, the reconstruction is reliable. When the pattern is weak, because the subject is rare, recent or too specific, the reconstruction produces something with the right shape and the wrong content. That is called hallucination, and it comes from the design, not from a manufacturing defect.

There is also a cut-off in time. Training ended on some date, and everything that happened after that simply is not in there. A model answering about recent events is searching the internet at that moment, not remembering.

How a model is made, in three phases

The first phase is pre-training. The system receives an immense volume of public text and is adjusted to predict the continuation. That phase is by far the most expensive of all: it is what consumes the months of computation on thousands of graphics cards the labs talk about, and it is what creates raw capability. At the end of it, the model knows a lot and does not know how to converse: it completes text, including completing your question with other similar questions.

The second phase is instruction tuning. The model is trained with examples of a request and an appropriate reply, and it learns that when somebody writes a question, the expected continuation is an answer, not more questions. That is the phase that turns a text completer into an assistant.

The third is alignment, in which people compare model answers and indicate which are better, and that judgment is used to adjust behaviour. This is where the tone, the refusals on certain subjects and the tendency to explain rather than dump come from. It is also where a known vice comes from: the model learns that confident answers please more than hesitant ones, and becomes confident even when it should not be.

More recent models add a reasoning step, in which the system is trained to produce an internal draft before the final answer. That greatly improves maths, logic and programming, and it charges you in time and in cost, because the draft is also generated and also billed.

Why it does not remember you

Every answer is calculated from scratch. The model keeps nothing between one call and the next. When it looks like it remembers the start of the conversation, what happened is that the program around it resent the whole conversation along with your new message.

That has two day-to-day consequences. The first is the limit: there is a ceiling on how much text fits in one call, called the context window, and long conversations eventually break that ceiling and start losing the beginning. The second is cost: resending the whole conversation with every message means paying for it again with every message. That detail is explained with the arithmetic in tokens and the context window.

The memory features some products offer do not change the mechanism: they keep notes about you outside the model and inject those notes at the start of each new conversation. It is a file on the outside, not a memory on the inside.

For anyone about to build something, that is the most useful piece of information in this article. Every behaviour of the system is defined by what goes into the call. If the information is not in the text you sent, it does not exist for the model, however obvious it seems to you.

What it does very well

There is a clear pattern in what works, and it is useful for deciding where to apply it. An LLM is excellent at transformation tasks: taking a text that exists and giving it back in another format, another tone, another language, another length.

Picture an accounting firm that receives thirty client messages a day on WhatsApp, each written differently. Turning those messages into standardized records, with subject, urgency and what the client is asking for, is exactly the kind of work the model gets right almost every time, because all the necessary information is already in the message. It does not need to know anything about the world, it needs to reorganize.

It is also very good at drafting. First version of a proposal, standard reply, product description, video script, difficult email. The gain is not having the final text, it is leaving the blank page in thirty seconds with something to edit.

And it is surprisingly good at classifying and extracting. Reading a hundred customer reviews and splitting them by type of complaint. Reading a contract and listing deadlines and amounts. Reading a CV and extracting experience. These are tasks where the material is all in the input and the output is structured.

What it does not do, by design

It is not a source of fact. Names, dates, numbers, amounts, statutes, bibliographic references: all of that is reconstructed from a pattern, not looked up. Where being wrong is expensive, the model text is a draft for checking, not a decision. That does not improve with a better model, it improves with the fact being handed over alongside the question.

It does not calculate reliably on its own. Reasoning models improved a great deal at this, but the base operation is still text prediction. For money, tax, payroll or anything where the arithmetic has to add up, the calculation belongs in code, and the model comes in to explain the result. At ROO3 that is a project rule: amounts are not calculated by a language model.

It has no intent and no opinion of its own. When it agrees with you after you insist, that is not persuasion: it is the effect of training that rewards agreeable answers. That is one of the most dangerous traps in professional use, because it looks like validation and is not.

And it does not know what it does not know. There is no internal confidence meter it can consult before answering. That is why the same assured voice comes out for what it got right and for what it made up.

What changes between one model and another

All large models perform the same operation, and there is still real difference between them. It shows up in four places: how much text fits in a call, how well the model does on hard reasoning and coding tests, how much it costs per volume of text, and how fast it answers.

Those four things move in opposite directions. A more capable model is usually more expensive and slower. A fast, cheap model handles classifying and extracting, and stumbles on long reasoning. The right choice depends on the task, and using the most expensive model for everything is the most common cost mistake in AI projects.

To compare without relying on opinion, there are two public measurements worth using: a composite capability index, which combines a model result across dozens of hard tests into a single scale, and a human preference score, in which people compare answers blind. They measure different things and frequently disagree.

Both measurements are gathered and explained in the AI Benchmark, which updates itself every day. ROO3 does not measure any model: the page gathers and credits public data from the people who do. If you want to know which one to use today, the subject is covered in which AI is best today.

What this changes in your company

The practical conclusion is a triage rule. Tasks where all the necessary material is already in the input and the result is quickly verifiable are the ones that pay off now, with little risk. Tasks where the model has to know something about the world on its own and nobody is going to check are the ones that generate losses.

Between those two extremes sits most real work, and for it there is a known solution: hand over the fact alongside the question. Instead of asking the model what your returns policy is, you send the policy with it and ask it to answer based on that. It has a name, it is called RAG, and it is the design most serious projects use.

The second rule is about where the human sits. Not reviewing everything, which cancels out the gain, but at the points where being wrong is expensive: before it goes to the customer, before it is written to the system, before anything involving money.

If you are deciding where to start in your company, the path in order is in AI for small business. And if you would rather have somebody look at your specific case before you invest, that is what ROO3 AI consulting does in the first stage.

Frequently asked questions

What does the acronym LLM mean?

LLM comes from large language model. Large refers to the size of the model, measured in billions of parameters, and to the volume of text used in training. It is the technology behind tools such as ChatGPT, Claude and Gemini.

Does an LLM understand what it is saying?

The debate is genuinely open and depends on what you call understanding. From the practical point of view, it behaves like something that understands across many tasks and fails in ways nothing that understood would fail, such as getting a date wrong inside a technically impeccable text.

Why does AI make things up so confidently?

Because it does not consult a bank of facts: it reconstructs the answer from statistical patterns. When the pattern is weak, the result comes out with the right shape and the wrong content. And training rewards confident answers, so there is no hesitation warning you that this is a fragile reconstruction.

Does the AI learn from my conversations?

The model itself does not change while you talk to it: it is fixed after training. Whether or not your conversations are used in future training depends on the policy of the product you use, and business plans normally exclude that by contract. It is worth reading the terms before putting customer data in.

What is the difference between an LLM and AI?

Artificial intelligence is the whole field, which includes quite different things such as computer vision, recommendation and forecasting. An LLM is a specific type of model, specialized in text. Every LLM is AI, but a large share of the AI running in the world today is not an LLM.

Do I need to understand this to use AI in my company?

You do not need the maths, but you do need the mechanism. Knowing that the model predicts text rather than looking up facts is what makes you hand over the information alongside the question instead of hoping it knows, and that is the difference between a project that works and one that generates rework.

Sources
Rodrigo Fávaro

Rodrigo Fávaro

Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.

X @rodmf LinkedIn rodrigofavaro

Want to apply this in your company?

ROO3 diagnoses what can be automated first in your business. The first conversation is free.