What agentic AI is: the difference between a chatbot that answers and an agent that acts
The word became fashionable in 2026 and almost lost its meaning. The technical difference is real, and it is what separates a project that saves time from one that creates losses.
Agentic AI is a system that receives an objective instead of a question, decides for itself which steps to take, uses tools to carry out each step, looks at the result and corrects the plan until it is done. The difference from a chatbot is not the intelligence of the model: it is the loop. The chatbot answers and stops. The agent acts, observes what happened and decides again, and that autonomy brings the gain and the risk at the same time.
What you get from this article
- What defines an agent is the loop of acting, observing and replanning, not the quality of the model.
- A chatbot takes a question and returns text; an agent takes an objective and returns work done.
- Errors in an agent compound: one wrong decision at step two contaminates every step after it.
- Autonomy should be proportional to the cost of a mistake, not to what the technology allows.
- The agent performs reversible actions; irreversible actions require human confirmation.
- Without a log of every step, an agent that goes wrong is impossible to audit.
The difference is in the loop, not the model
A chatbot works in a straight line: you ask, it answers, done. If the answer is wrong, you ask again. The cycle is always started by you, and the system never does anything on its own.
An agent works in a circle. It receives an objective, plans the steps, executes the first one using some tool, looks at what happened, decides whether the plan still holds, adjusts if it does not, executes the next. It repeats until it finishes or gives up.
That phrase in the middle, looking at the result, is the whole thing. It is what lets the system notice that the search returned nothing and try another term, that the file was in another format and convert it first, that the customer is already on file and update rather than duplicate. Without that step, you have an automated sequence, which is useful and is not an agent.
And here is the counterintuitive part: the same model can be both. What changes is the program around it. An agent is not a type of model, it is a way of organizing the use of one.
Nobody buys an agent. You build the loop around a model the company already pays for.
The four pieces an agent needs
An objective. Unlike a question, an objective describes a desired end state without saying how to get there. "Answer this question" is a question. "Organize today quote requests, separate the ones missing information and ask for what is missing" is an objective.
Tools. With no reach into the world, the agent only thinks. It has to be able to search, read, write, call a system, send. This is where MCP comes in, which standardized that reach and is why agents became viable outside the lab.
Working memory. The agent has to know what it already tried, what worked and what failed. Without that, it repeats the same step in a loop, which is the most common and most expensive failure mode, because every repetition costs money.
A stopping rule. When to consider it done, how many attempts before giving up, when to call a person. An agent with no stopping rule is an invoice generator, and that is not a turn of phrase: it is the most common way an AI project blows its budget in one night.
Where this works well today
It is worth being specific, because the distance between the demo and real use is large and the market promise is far too generous.
Agents work well when the result is verifiable and the environment is closed. Programming is the clearest example: the agent writes code, runs the test, sees the error, fixes it, runs it again. The test is an objective judge, and the loop has somewhere to converge. That is why the most mature category of agent today is the coding assistant, such as Claude Code.
It also works well in research and gathering: finding information across several sources, comparing it, assembling a summary with references. The error here is visible and the cost of being wrong is low, because a person reads it before using it.
And it works in multi-step triage: reading the day messages, classifying them, checking in the system whether that customer exists, drafting the reply for the simple cases and separating the complex ones for a person. It is routine work, with clear rules and a checkable result. Triage, customer service and follow-up are neighbouring fronts, and they are described side by side in AI for business.
It still works badly on long, open-ended tasks in a changing environment. The more steps, the more chance the agent goes off the rails, and that chance does not grow in a straight line: it compounds.
The error that compounds, and why it is frightening
This is the structural problem of the approach, and understanding it is what separates careful rollouts from painful ones.
In a chatbot, an error is an error: the answer came out badly, you notice and ask again. In an agent, the error at step two becomes the input to step three. The system carries on working from a wrong premise, with all the confidence in the world, and the final result can be completely off with no warning sign along the way.
Do the arithmetic with easy numbers. If each step has a 95% chance of coming out right, which is pretty good, a twenty-step task has about a 36% chance of finishing error-free. A sequence of very reliable steps produces an unreliable result. That is not pessimism, it is multiplication.
The design consequence is direct and it is the most important thing in this article: prefer short agents chained by you over one long autonomous agent. Five four-step tasks, each with a check at the end, is far more reliable than one twenty-step task. And when it fails, you know exactly where.
Autonomy proportional to the cost of a mistake
The question that decides the design is not what the technology can do. It is what happens if it goes wrong, and who pays.
There is a ladder of autonomy, and it is simple to apply. On the lowest rung, the agent only suggests and a person executes. On the next, it executes and the person confirms first. Then it executes on its own and reports afterwards. At the top, it executes and nobody looks, except in a report.
The practical rule that avoids nearly every problem: the agent performs reversible actions on its own; irreversible actions require authorization. Creating a draft, classifying, marking as read, assembling a proposal, all reversible. Sending to the customer, deleting, issuing an invoice, paying, cancelling an order, not.
The temptation to jump straight to the top is strong, because that is where the most visible time saving sits. But the saving is asymmetric: you gain minutes per run and can lose a customer in one. Climbing one rung at a time, after weeks of logs showing that the previous rung is reliable, is slow and is what works.
- Suggests and a person executes: zero risk, real time saving already.
- Executes with confirmation: good for what is repetitive and visible.
- Executes and reports: only after the log shows weeks without error.
- Executes unsupervised: only for reversible, cheap, low-impact actions.
What has to exist before switching it on
A log of everything. Every step, every tool called, every result, every decision. Without it, when the agent does something strange, nobody will be able to reconstruct what happened, and the system becomes a box that sometimes goes wrong with no explanation. That log is the item people most regret not having.
Cost and attempt ceilings. Maximum steps per task, maximum calls per hour, maximum spend per day. An agent stuck in a loop does not announce that it is in a loop: it looks busy. The ceiling is what turns a fright into a small incident.
Minimum permissions. Each tool with the smallest credential that does the job. Read separated from write. No single administrative credential because it was easier to configure.
Distrust of what comes back. Content returned by a tool can contain text addressed to the model, trying to redirect its behaviour. That is not a theoretical hypothesis: it is the most studied attack vector in this kind of system. A tool result is data, never instructions.
Where a company starts
The honest recommendation is to start with what looks unambitious, because that is what survives the first month.
Choose a task your team does several times a day, with clear rules, whose result can be checked in seconds and whose mistakes are cheap. Message triage, filling in a form from loose text, a first quote reply using data already in the system.
Run it in suggestion mode for two or three weeks, with logging. Measure two things: how often the suggestion was accepted unchanged and how much time it saved. If the acceptance rate is high and stable, climb a rung. If it is unstable, the problem is not autonomy, it is the task or the instruction.
The most common mistake is starting with customer service at full autonomy, because that is what shows up in demos. It is precisely the worst place to start: the mistake is public, the cost is the customer relationship and the variety of situations is at its maximum.
If you want to decide which task in your operation is the right candidate and which rung of autonomy it can take, that is the diagnosis that opens an AI consulting engagement at ROO3.
Frequently asked questions
What is the difference between a chatbot and an AI agent?
A chatbot takes a question and returns text, always started by you. An agent takes an objective, decides the steps, uses tools, looks at the result of each step and replans until it is done. The difference is in the act-and-observe loop, not in the quality of the model.
Do I need a special model to build an agent?
No. The same model can work as a chatbot or as an agent, depending on the program around it. What changes is the architecture: available tools, memory of what has been tried, and a stopping rule.
Why do agents fail on long tasks?
Because errors compound. If each step has a 95% chance of being right, a twenty-step task has about a 36% chance of finishing error-free. That is why several short agents, with checks between them, are more reliable than one long autonomous agent.
Is it safe to give an agent access to my systems?
It depends on the design. The rule that avoids most problems is autonomy proportional to the cost of a mistake: the agent performs reversible actions on its own, and irreversible actions such as sending, deleting or paying require human confirmation. Add minimum permissions per tool and a log of every step.
Does an AI agent replace an employee?
In practice it absorbs routine tasks with clear rules and verifiable results, and it still needs people at the judgment calls and the exceptions. Where companies go wrong is sizing the saving as if supervision cost nothing, and it does not.
What does running an agent cost?
More than an equivalent chatbot, because every step is a call and the task history travels with all of them. That is why step and daily spend ceilings are not a configuration detail: they are what separates predictable cost from losses.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

What MCP is, the standard that connects AI to your tools
What the Model Context Protocol is, the problem it solves, how it works in practice, who adopted it, and the security...
10 min read
What Claude Code is and how to use the Anthropic agent
What Claude Code is, how it differs from autocomplete, what it can do on its own, and the precautions before giving it...
9 min read
AI for small business: where to start without wasting money
The practical path for a small business to start using AI: how to choose the first task, what to measure, what it costs...
11 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.