Which AI is best today: how to compare properly instead of following internet opinion
The question does not have one answer, it has two. And they disagree with each other often enough that it should bother more people.
There is no single best AI, there are two different questions with different answers. For capability measured on hard tests, the most solid index is the Epoch AI ECI, which combines performance across more than fifty benchmarks into a single scale. For the preference of ordinary people in everyday conversation, the Arena ELO applies, where people choose blind between two answers. The two rankings frequently disagree, and that disagreement is information, not error.
What you get from this article
- Capability and human preference are two different measures, and the top of each rarely coincides.
- The Epoch AI ECI combines more than 50 benchmarks into a single scale, anchored to allow comparison over time.
- The Arena ELO measures what people prefer to receive, which is not the same as raw capability.
- Neither measures what decides your case: price, speed and quality in your language.
- The test that settles it is running your ten real tasks on the candidates, with the answer key written in advance.
- The lead changes every few months: build an architecture that lets you switch models.
Why the question has more than one answer
Asking which AI is best is like asking which vehicle is best. The answer changes completely depending on whether you are carrying seven people, crossing the city at rush hour, or hauling five hundred kilos of sand.
With models, there are at least two questions hidden inside one. The first is about capability: which one solves the hardest problems, the ones requiring long reasoning, maths, programming and specialist knowledge. The second is about preference: which one people most like receiving as an answer in an ordinary conversation.
Both are legitimate and they measure different things. A model can solve an olympiad maths problem and write in a way nobody enjoys reading. Another can be pleasant, direct and well calibrated in tone, and stumble on a problem that needs fifteen steps of reasoning.
That is why serious rankings do not collapse everything into a single score. When somebody publishes one number, they chose for you which of the two questions matters, and they probably did not say so.
A single score is comfortable to read and impossible to audit. Two columns that disagree with each other say far more.
The ECI: capability measured on hard tests
Epoch AI is a research organization that tracks the technical evolution of models, and the index it publishes, the Epoch Capabilities Index, solves an awkward comparison problem.
The problem is this: benchmarks saturate. A test where every model scores top marks stops distinguishing anything, and gets replaced by a harder one. Except then you can no longer compare today model with the one from two years ago, because they took different tests.
The ECI solves that by combining performance across more than fifty distinct benchmarks into a single scale, using a statistical method that estimates at the same time the difficulty of each test and the capability of each model. It is the same family of technique used to calibrate exams in education, and it allows comparison of models that never took exactly the same tests.
To give a sense of the scale, it is anchored at known points: Claude 3.5 Sonnet was fixed at 130 and GPT-5 at 150. Higher numbers mean more capability, and the distance between them has a consistent meaning over time, which is exactly what an isolated benchmark cannot offer.
The Arena ELO: what people prefer
The other measurement works on a completely different principle, and it is easier to explain: people send the same question to two models without knowing which they are, read both answers and pick the better one.
With many votes, that becomes a score calculated on the same rating scheme used in chess, the ELO. Beating a well-placed model is worth more than beating a poorly placed one, and the score adjusts continuously.
What that measurement captures is real and matters: clarity, appropriate tone, an answer of the right length, not padding, understanding what the person actually wanted. These are qualities no technical benchmark measures and that define the experience of anybody using the tool every day.
And it has a known bias, worth knowing so you do not over-read it: because the voters are people asking ordinary questions, the ranking rewards what pleases in everyday conversation. A well-formatted, friendly, confident answer tends to win. That favours models tuned to please, and pleasing is not always being right.
When the two disagree, and why that is useful
Disagreement between the two lists is not a flaw in either: it is the most informative thing in the set.
A model that is high on capability and low on preference is usually the one that solves the hard problem and answers in a tiring way: too long, too technical, not adjusted to what the person asked. For an automated system, where nobody reads the raw answer, that matters little. For an assistant talking to a customer, it matters a lot.
A model that is high on preference and lower on capability is pleasant, direct and well calibrated, and will stumble when the task requires long reasoning. Great for customer service and copywriting. Risky for complex analysis.
The practical reading, then, is this: look at capability when the task is hard and at preference when a person will read the text. If both point at the same model, the choice is easy. When they point at different models, they have just told you that you have two tasks, not one.
Both measurements, with credit to the original sources, are gathered in the AI Benchmark, which updates itself every day. ROO3 does not measure any model and is not affiliated with the organizations that do: the page gathers, translates and explains public data.
What no ranking measures, and what decides your choice
Here is the part that almost never shows up in the which-is-best debate, and that in practice decides more than the score does.
Price. The cost difference between the most capable model and a mid-tier one can be several times over. If the task is classifying messages, paying for the top of the ranking is pure waste. The arithmetic of how that scales is in tokens and the context window.
Speed. A model that thinks more answers more slowly. In an overnight report, that is irrelevant. In a chat with a customer waiting, eight seconds of silence is a lost conversation.
Quality in your language. The rankings are dominated by English-language evaluation. A model that is excellent there can write correct but unnatural Portuguese, with constructions that sound translated. That only shows up by testing with your own material.
Access to what just happened. Both rankings evaluate closed questions, and neither asks whether the model knows what was in the news this morning. For anyone working with breaking subjects or public reaction, that eliminates candidates without moving a single score.
Integration and data policy. Does the model connect to what you already use? Does the contract allow sending customer data? Is there a commitment not to use your content in training? These questions eliminate candidates regardless of ranking position.
One example of how a criterion like that decides on its own: Grok reads X in real time, and that moves not one point in either ranking. For anyone who needs it, the candidate list shrinks to one. For anyone who does not, it is irrelevant.
- Price per volume of text, counting input, output and reasoning.
- Response time, when somebody is waiting on the other end.
- Naturalness in your language, tested with your own material.
- Access to recent information, which depends on searching live and not on what the model trained on.
- Integration with the systems the company already uses.
- Data policy and the existence of a business plan with a contractual commitment.
The test that settles it in one afternoon
After all of that, the method that actually decides is embarrassingly simple, and almost nobody does it.
Write down ten real tasks of yours. Not demo questions: the things you would genuinely do, with your material, in your vocabulary. An email you need to write, a document you need to summarize, a customer message you need to classify.
Write the answer key first. What a good answer to each would be. That looks bureaucratic and it is the step that separates evaluation from impression: without criteria written in advance, you judge by the first answer that sounded nice.
Run the ten on two or three candidates and compare. In one afternoon you have an answer based on your case, worth more than any general ranking. And keep those ten tasks: they become your standard test every time a new model ships.
One last recommendation worth more than the choice itself: build in a way that lets you switch. The lead changes every few months, and a system tied to one specific vendor turns every switch into a project. Open standards such as MCP exist in large part because of this. If you want to set that test up with the right rigour for your case, that is how ROO3 AI consulting begins.
Frequently asked questions
Which AI is best today?
It depends what you call best. For capability measured on hard tests, look at the top of the Epoch AI ECI. For the preference of people in ordinary conversation, look at the top of the Arena ELO. The two rankings frequently disagree, and the lead changes every few months.
What is the Epoch AI ECI?
It is the Epoch Capabilities Index, a composite index that combines a model performance across more than fifty benchmarks into a single scale. It uses a statistical method that estimates at the same time the difficulty of each test and the capability of each model, which allows comparing models that took different tests.
What is the Arena ELO?
It is a human preference score. People send the same question to two models without knowing which they are, pick the better answer, and the system calculates a score on the same rating scheme as chess. It measures what people like to receive, which is not the same as raw capability.
Why do the two rankings disagree?
Because they measure different things. A model can solve very hard problems and answer in a long, tiring way, while another is pleasant and direct but stumbles on long reasoning. The disagreement is useful information: it indicates which model suits a technical task and which suits text a person will read.
Do I always have to use the model at the top of the ranking?
No, and using the most capable one for everything is the most common cost mistake in companies. For classifying, extracting and labelling, a small model delivers the same result for a fraction of the price. The top tier is justified for text that goes to the customer and for tasks with many reasoning steps.
How do I test which model is best for my case?
Write down ten real tasks of yours, with the answer key for what a good response would be defined before testing, and run the ten on two or three candidates. In one afternoon you have an answer based on your material, and those ten tasks become your standard test for every new model that ships.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

ChatGPT, Claude or Gemini: how to choose for your use
A practical comparison of ChatGPT, Claude and Gemini: what each does better by design, what does not change with every...
9 min read
What an LLM is: how it works inside and what it does not do
What an LLM is, explained without maths: how the model predicts the next word, why that works so well, and which limits...
12 min read
Tokens and the context window: why the AI bill comes in high
What a token is, what a context window is and how billing really works. With the arithmetic done and the four mistakes...
11 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.