Fundamentals

What a token and a context window are, and why the AI API bill is bigger than you calculated

Almost every budget overrun in an AI project has the same origin: somebody calculated the cost of one question and forgot that the whole conversation travels along with all the others.

Rodrigo Fávaro, fundador da ROO3
Rodrigo Fávaro Founder of ROO3
·11 min read
Abstract illustration of text blocks piling up with each round of conversation, the stack growing in neon green on a black background.
Short answer

A token is the chunk of text the model processes, roughly three quarters of a word. The context window is how much text fits in a single call, counting what you send and what it answers. The bill surprises people because the model has no memory: with every new message, the whole conversation is resent and billed again, which makes the cost of a conversation grow quadratically rather than in a straight line.

What you get from this article

  • A token is about 0.75 of a word, and accented characters and proper nouns consume more than expected.
  • The context window is the ceiling of one call, not the model memory between calls.
  • Conversation cost grows quadratically: the tenth message pays for the previous nine again.
  • Output costs several times more than input, and reasoning models bill the draft you never see.
  • Prompt caching cuts the cost of the fixed block and is the highest-return adjustment in almost any project.
  • A ceiling per user and per day is not a detail: it is what separates a predictable cost from a loss.

Tokens: the unit the machine actually reads

The model does not read letters or words. It reads tokens, chunks of text of variable length defined by a dictionary created during training. Common words are usually one whole token. Rare words break into several. Punctuation, spaces and line breaks count too.

The most useful rule of thumb: one token equals roughly 0.75 of a word. A thousand words comes to around 1,300 tokens. A page of continuous text runs around 600 to 800 tokens. A ten-page contract easily passes 7,000.

There is a specific trap for languages other than English that almost nobody factors into a budget. The dictionaries of these models were built mostly on English text, so accented words, local proper nouns, addresses and sector-specific terms tend to fragment into more tokens than the English equivalent. In practice, the same content costs more in Portuguese than it would in English. Do not leave here with a percentage: measure your own text with the provider token counter, because the difference changes with your sector vocabulary.

Add to that what is not visible text: system instructions, descriptions of available tools, JSON formatting, markup. All of that is tokens, and all of it is billed every time.

Context window: the ceiling of one call

The context window is the total number of tokens that fits in a single call, counting everything you send and everything the model answers. If the window is 200,000 tokens, that is what has to hold the instruction, the history, the attached documents and the answer.

The most frequent confusion is thinking a large window means memory. It does not. The model keeps nothing between one call and the next: a large window means more fits at once, and nothing else. What resends the conversation with every message is the program around it, not the model.

When the conversation passes the ceiling, something has to go, and that is where the behaviour that unsettles users appears: the assistant forgets what was agreed at the start. It is not a failure, it is the cut. Tools handle it in different ways, summarizing the beginning or discarding it, and each choice has a cost in fidelity.

There is also a measured and little-publicized effect: models tend to pay more attention to the beginning and the end of what was sent than to the middle. Stuffing a giant document into the window and hoping that a sentence lost on page 40 gets used is a worse bet than it looks. Cutting out the relevant passage and sending only that usually gives a better and cheaper result.

A large window is not memory. It is just a bigger truck, still loaded from empty on every trip.

The arithmetic that surprises: why cost grows quadratically

Here is the point that blows budgets, and it is simple to demonstrate. Because the model does not remember, every new message has to carry the whole previous conversation with it. So you do not pay once per message: you pay for the accumulated conversation, again, every round.

Let us do the arithmetic with round, illustrative numbers. Suppose a support conversation where the system instruction is 1,000 tokens, each user question is 100 and each answer is 300. In the first round, 1,100 input tokens go in, the 1,000 of the instruction plus the 100 of the question, and 300 come out. In the second, the same 1,000 of instruction go in, plus the 400 the previous round left in the history, plus the 100 of the new question: 1,500. In the third, 1,900. Each round adds a fixed 400, and the tenth goes in at 4,700.

Adding up the ten rounds, you did not process 11,000 input tokens, you processed 29,000. Almost triple what a naive calculation would estimate, and the difference grows the longer the conversation, because the history goes in again every round. Multiply by a thousand conversations a month and the estimation error becomes the entire budget.

The second surprise is the price asymmetry. In the public price lists of the main providers, an output token typically costs five times an input token, and more than that in some models. A long answer costs disproportionately much, and asking the model to be concise is a cost measure, not just a style one.

The invisible draft of reasoning models

Models with reasoning produce an internal draft before answering. That draft is generated, counted as output and billed, even when you do not see it on screen.

That creates two practical traps in a project. The first is cost: a 200-token answer can have cost 2,000 output tokens, and the difference shows up only on the invoice, not in what you read.

The second is technical and breaks integrations in production. If you set a ceiling on output tokens, the draft shares that ceiling with the final answer. A long piece of reasoning can consume the whole limit and return an empty or truncated answer. The program expecting text receives nothing, and the error looks random because it only happens on the hard questions.

Where the task is simple and repetitive, such as classifying a message or extracting a field from a document, deep reasoning does not improve the result and multiplies the cost. Reducing the reasoning effort in those cases is one of the highest-return adjustments there is.

Prompt caching: the highest-return adjustment

In almost every AI system there is a fixed piece that goes in every call: the system instruction, the tone guide, the product list, the company rules. That block does not change and is resent thousands of times.

Prompt caching exists exactly for that. The provider keeps the processing of that fixed passage for a while and charges the next read at a fraction of the normal input price. Writing to the cache costs a little more than an ordinary input; reading from it costs much less.

The golden rule is architectural, and it is what most people get wrong: what is fixed has to come first and never change, and what varies comes after. If you put today date or the customer name at the start of the instruction, the fixed passage changes with every call, the cache never hits and you pay full price without understanding why.

In systems with a long instruction and many calls, that adjustment alone usually cuts a good part of the bill. It is the first thing to look at when an AI project gets too expensive, before switching models.

Choose the model by task, not by reputation

The most common cost mistake in companies is using the most capable model for everything. It is expensive, it is slower and, on most routine tasks, it does not deliver a better result.

A simple triage solves it. Classifying a message, extracting a field, labelling, deciding which department to route to: a small, cheap model with reasoning at minimum. Writing text that goes to the customer, analysing a document, solving a multi-step problem: a capable model. The saving from running the routine on the small model usually funds the use of the expensive one where it matters.

There is a pattern that works well and is underused: the cheap model does the triage and only escalates to the expensive one when the task is genuinely hard or when confidence is low. Most of the volume never reaches the expensive model.

To compare capability without relying on marketing, it is worth looking at the public measurements gathered in the AI Benchmark. And before switching models over cost, do the arithmetic on this page: frequently the problem is not the model price, it is the architecture of the call.

The four mistakes that multiply the bill

First: dragging the whole conversation forever. Past a certain point, summarizing the history and sending the summary costs a fraction of resending everything, and the result barely worsens. Almost no home-grown system does this.

Second: sending the whole document. Attaching an 80-page PDF to answer a question that is in paragraph 3 is paying for 80 pages to get one. Finding the right passage and sending only that is cheaper and gives a better answer, because the model does not get lost in the middle.

Third: having no ceiling. Without a limit per user, per session and per day, any badly closed loop or abusive use becomes an invoice. That is not hypothetical: a badly written scheduled job repeating a call is the fastest known way to spend a month budget in one night.

Fourth: not measuring. If you do not record input, output, reasoning and cache tokens per call, you have no way of knowing where the money went. Writing that to a table from day one costs almost nothing and is the difference between optimizing with data and optimizing with guesswork.

How to budget a project without fooling yourself

The method that works has four steps and takes an afternoon. First, write one complete real interaction, the way it will actually happen, with the real system instruction rather than a shortened version.

Second, count the tokens of that interaction with the provider token counter, not by eye. Third, multiply by the number of interactions you expect per month and add the accumulated effect of the conversation, that quadratic growth from the third section. Fourth, and this is the step almost everybody skips, multiply the total by three.

The factor of three is not pessimism, it is what experience shows: retries when something errors, conversations longer than predicted, people testing, the reasoning draft you forgot to count. If the project only works on the optimistic estimate, it does not work.

And there is the rule ROO3 applies to its own products: the selling price of any AI feature needs enough margin to absorb twice the estimated cost. Anyone selling unlimited with no ceiling is betting that the customer will use it lightly. If you are designing a product like that now, that conversation is part of what AI consulting resolves before it becomes code.

Frequently asked questions

How many tokens are there in a word?

On average a token equals about 0.75 of a word, so a thousand words comes to around 1,300 tokens. Texts with many accented characters, local proper nouns and technical terms fragment more and consume more tokens than the English equivalent, so it is worth measuring your own material in the provider token counter instead of applying a rule-of-thumb percentage.

Does a large context window mean the AI remembers me?

No. The window is the ceiling of a single call, not memory between calls. The model is recalculated from scratch every time, and when it seems to remember it is because the program resent the whole conversation with the new message, paying for it again.

Why was my API bill bigger than I calculated?

Almost always for three reasons combined: the whole conversation is resent with every message, which makes the cost grow quadratically; output typically costs five times more than input; and reasoning models bill the internal draft you never see on screen.

What is prompt caching and is it worth it?

It is keeping the processing of the fixed part of your instruction so that subsequent calls read it at a fraction of the price. It is very worthwhile in systems with a long instruction and many calls. The condition is that the fixed block comes first and does not change, otherwise the cache never hits.

Should I always use the most advanced model?

No. For classifying, extracting and labelling, a small model delivers the same result for a fraction of the price and with less waiting. Reserve the capable model for text that goes to the customer and for tasks with several reasoning steps.

How do I estimate the cost before building?

Write one complete real interaction, count the tokens with the provider tool, multiply by the expected volume adding the accumulated growth of the conversation, and then multiply the total by three. The factor of three covers retries, heavier use than predicted and the reasoning draft.

Sources
Rodrigo Fávaro

Rodrigo Fávaro

Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.

X @rodmf LinkedIn rodrigofavaro

Want to apply this in your company?

ROO3 diagnoses what can be automated first in your business. The first conversation is free.