Tools

Opus 5.5 is out: it wins seven of nine tests and costs 60% less than Fable 5.1 on output

Anthropic published the comparison against Fable 5.1, Opus 5 and GPT-6 Astra. The number that actually decides your bill is not in the headline: it is the cache read price, now down to US$ 0.20.

Rodrigo Fávaro, fundador da ROO3
Rodrigo Fávaro Founder of ROO3
·11 min read
Comparison table showing Claude Opus 5.5 against Fable 5.1, Opus 5 and GPT-6 Astra across nine performance tests and on price per token.
Short answer

Anthropic shipped Claude Opus 5.5 on 22 September 2026 at US$ 4 per million input tokens and US$ 20 per million output, 20% below Opus 5 and 60% below Fable 5.1. In the comparison the company published itself, against Fable 5.1, Opus 5 and GPT-6 Astra, it wins seven of the nine tests, including all three coding ones. It loses two to GPT-6 Astra: automation and terminal science. The number the headline does not highlight is the cache read price, down from US$ 0.50 to US$ 0.20 per million, a 60% cut that matters more than the input price for anyone running an agent or a chat with a long system prompt.

What you get from this article

  • The model is already in Anthropic official catalog as claude-opus-5-5.
  • Price: US$ 4 per million input tokens and US$ 20 per million output. Fable 5.1 charges US$ 10 and US$ 50.
  • Cache reads fell from US$ 0.50 to US$ 0.20 per million, a 60% cut. That is the line that moves a real bill the most.
  • It wins 7 of the 9 published tests, including all three coding benchmarks.
  • It loses 2 to GPT-6 Astra: AutomationBench (40.0% against 41.4%) and Terminal-Bench-Science (58.7% against 64.6%).
  • It generates output more than 30% faster than Opus 5, and Anthropic says it runs 40% cheaper on typical workloads.
  • Fast mode doubles the sheet: US$ 8 input and US$ 40 output.
  • Sonnet 5.5 and Haiku 5.5 were promised for the coming weeks.

What shipped, and how to confirm it without trusting a screenshot

Claude Opus 5.5 shipped on 22 September 2026, without the weeks of postponed announcements that marked the Grok 4.7 launch. It went straight from the official page to the catalog.

The way to confirm it is the same as always, and it is not the screenshot going around: the identifier claude-opus-5-5 is in Anthropic documentation, with published pricing, and the model is already on the Claude Platform, AWS, Google Cloud and Azure. Anyone can mock up a table image.

Alongside the model came the official comparison against Claude Fable 5.1 and Claude Opus 5, both in-house, and against GPT-6 Astra, from OpenAI. That is the table reproduced below, with the figures checked against the announcement.

The table, with nine tests and three price lines

Two things to read first. One: GPT-6 Astra prices were not in Anthropic table, so its column is empty on the price rows. We prefer an empty cell to a guessed number.

Two: four of the nine rows measure coding or long autonomous terminal work. That is the fight Anthropic chose to pick, and it is where the gap comes out widest.

Opus 5.5 against Fable 5.1, Opus 5 and GPT-6 Astra, in Anthropic own figures
Opus 5.5 new Fable 5.1 Opus 5 GPT-6 Astra
Input price US$ per million tokens US$ 4 US$ 10 US$ 5 not published
Output price US$ per million tokens US$ 20 US$ 50 US$ 25 not published
Cache read US$ per million tokens US$ 0.20 not published US$ 0.50 not published
Long terminal work Terminal-Bench 4.0 66.4% 55.8% 52.3% 57.9%
Coding FrontierCode v1.1 54.4% 50.3% 48.0% 53.3%
Coding CursorBench 4.0 57.8% 51.8% 46.6% not measured
Knowledge work GDPval-AA v2.1 1,846 1,735 1,708 1,542
Automation AutomationBench 40.0% 31.4% 26.9% 41.4%
Hard knowledge Humanity's Last Exam 67.7% 65.6% 63.6% 57.2%
Terminal science Terminal-Bench-Science 0.1 58.7% 52.6% 29.0% 64.6%
Computer use OSWorld 2.0 81.8% 80.7% 74.0% not measured
Chart reading Chartography 89.0% 88.4% 83.4% not measured
Figures published by Anthropic in the Opus 5.5 announcement. Opus 5 pricing is stated indirectly, through the announcement saying the new sheet is 20% lower. GPT-6 Astra prices and some of its scores are absent from the original table, so they appear empty here rather than estimated.

What flipped since last week

Yesterday we published the Grok 4.7 math here, and the question it left was blunt: are twenty points of Terminal-Bench worth paying eight times more on output? The cheap model won two tests of seven, and the expensive one won wherever the work was long and autonomous.

Opus 5.5 moves both sides of that equation at once, which is what makes it unusual. It went up on the test the expensive model already won, from Fable 55.8% to 66.4%, and it came down on price, from US$ 50 to US$ 20 on output.

In practice the distance between the top and the cheap end shrank from below. Grok 4.7 still costs US$ 6 on output against US$ 20 for Opus 5.5. That is still more than three times, but it is no longer the eight times of yesterday. And on the other side of the table, the capability gap on terminal work widened: 38.0% for Grok against 66.4% for Opus 5.5.

Anyone who was postponing a decision waiting for top-tier prices to fall now has a concrete reason to redo the spreadsheet. Anyone who picked the cheap model because of output cost should redo it too, because the number that justified the choice has changed.

It is rare for a launch to move both columns in the same direction. Normally you pay more for better, or accept less for cheap. Here the better one got cheaper, and that reopens decisions that were already closed.

Where it does not win, said plainly

Of the nine tests, GPT-6 Astra wins two, and Anthropic shows that in its own table. Recording it is not a detail: it is what separates a comparison from an ad.

On AutomationBench, which measures task automation, it is 41.4% for GPT-6 Astra against 40.0% for Opus 5.5. The gap is 1.4 points, well within what measurement variance would explain. Calling that a defeat is close to a stretch, but calling it an Opus win would be a lie.

On Terminal-Bench-Science, which measures scientific work inside a terminal, the gap is real: 64.6% for GPT-6 Astra against 58.7% for Opus 5.5. That is nearly six points, and it is the only row where the new model sits clearly behind someone.

That row is worth reading from another angle too. Opus 5 scored 29.0% on it. Going to 58.7% is a doubling in one generation, the largest internal jump in the whole table. The model improved enormously at what used to be its weakest point, and it is still behind the competitor. That is exactly what happened to Grok 4.7 on Terminal-Bench, and the reading is the same: a big jump is not the same as a lead.

The number that decides the bill is not in the headline

The launch headline is the input and output price. The line that will actually move your invoice is a different one: cache reads fell from US$ 0.50 to US$ 0.20 per million tokens, a 60% cut.

Cache, here, is the chunk of text that repeats on every call: the system prompt, the product catalog, the company handbook, the conversation history. You pay a low rate for the model to reread what it has already seen, instead of paying full input price every time.

In anything that actually runs, that chunk dominates the bill. A support chat with an 8,000-token system prompt, answering 3,000 conversations a month at 6 turns each, rereads that prompt 18,000 times. That is 144 million cache tokens. At US$ 0.50 it cost US$ 72. At US$ 0.20 it costs US$ 28.80. Input and output have not even entered that calculation.

This is why comparing models on input price alone misleads. Two models with identical input sheets can produce very different invoices depending on what they charge to reread what repeats, and almost nobody checks that line before deciding.

Headline price is about the new token. Real price is about the repeated one. In an application that runs every day, the repeated one is most of it.

The fine print: fast mode doubles the sheet

There is a line that was not in the launch image and that deserves the same attention as Grok 200,000-token cliff.

Opus 5.5 has a fast mode, and it carries its own sheet: US$ 8 input and US$ 40 output, exactly double the normal price. Anyone who switches fast mode on thinking they are only buying speed will find the price on the invoice, not at the moment of deciding.

Doubled, it costs almost what Fable 5.1 costs in normal mode. So the practical question is whether the extra speed is worth double. For an agent running alone overnight, almost certainly not: nobody is waiting. For a chat where a person is watching the screen while the answer streams, possibly yes, because there the wait has an abandonment cost.

Anthropic says the model already generates output more than 30% faster than Opus 5 in normal mode, and that it runs 40% cheaper on typical workloads, counting saved steps and not just the price sheet. That second number is the hardest to verify from outside, because it depends on what counts as typical. Treat it as a vendor estimate, not a measurement.

The safety part, and why it belongs here

Anthropic highlighted that Opus 5.5 attempted to circumvent boundaries about 85% less often than Opus 5, that it scored the best of any of its models on an automated behavioural audit, and that it is less likely to take hard-to-reverse actions.

This tends to get read as a lab concern, and it is not. Anyone putting a model to work unsupervised on a production system, deleting records, sending email on the company behalf or approving payments, is exposed to exactly this. A model that avoids irreversible action when unsure is worth more, in that setting, than a few extra points on a test.

The usual caveat: the party measuring is the party selling. The audit is Anthropic own, on Anthropic own model. The number works as a direction signal, not a guarantee, and it does not replace limiting what the model is allowed to do in your system. Tight permissions remain more reliable than good behaviour.

What changes for people using AI at work

If you use AI through the interface, on a monthly plan, almost nothing changes this week beyond the model being better and faster. You do not pick models by token price.

If you built something that calls the API, the table has a direct consequence, and it fits into five decisions.

What cannot be said yet

Two things stay open, and it is honest to list them instead of pretending the table settles everything.

First: there is no independent measurement of Opus 5.5 yet. Epoch AI, the source behind our AI Benchmark, has not published an ECI for it, which is normal: the model shipped today. It should show up there shortly in the recently launched band, with a dash where the score goes, which is how Grok 4.7 appeared on its own launch day. Until the ECI lands, all that exists is the vendor table. It is not false, but the vendor chose which nine tests to show.

Second: Sonnet 5.5 and Haiku 5.5 were promised for the coming weeks. Those set the price of the tier most companies actually use. If the Opus price cut repeats in the smaller models, the bill for anyone running volume changes far more than it changed today.

Frequently asked questions

Is Opus 5.5 better than Fable 5.1 and GPT-6?

In the comparison Anthropic published, Opus 5.5 wins seven of nine tests, including all three coding benchmarks and long terminal work. GPT-6 Astra wins two: automation, by 1.4 points, and terminal science, by nearly six. Fable 5.1 wins none, and costs two and a half times more on output.

How much does Opus 5.5 cost?

US$ 4 per million input tokens and US$ 20 per million output. Cache reads cost US$ 0.20 and cache writes US$ 5 per million. Fast mode doubles input and output, to US$ 8 and US$ 40. For comparison: Fable 5.1 charges US$ 10 and US$ 50, and Opus 5 charged US$ 5 and US$ 25.

Why does the cache price matter so much?

Because in an application that actually runs, most tokens repeat: the system prompt, the catalog, the conversation history. You pay the cache read rate every time the model rereads that chunk. In a chat with an 8,000-token prompt and 18,000 turns a month, the cut from US$ 0.50 to US$ 0.20 takes that part of the invoice from US$ 72 to US$ 28.80, before input and output.

Is it worth switching from a cheap model to Opus 5.5 now?

It depends what you run. For classifying messages, summarising and extracting data at volume, the cheap model still handles it and the quality gap does not show. For an agent working unsupervised for hours on code or systems, the capability gap widened and the price fell, so the math that came out negative last week may come out positive now. The right path is running your own tasks on both before changing anything in production.

Can these numbers be trusted?

As a starting point, yes, with the usual caveat: the table was published by Anthropic about an Anthropic model, and Anthropic chose which nine tests to show. It is worth noting that the table does show two losses to GPT-6 Astra, which is a sign of a less airbrushed comparison than average. Independent measurement, such as Epoch AI, does not exist for this model yet.

When does Opus 5.5 enter the ROO3 AI Benchmark?

In two stages, and neither uses vendor figures. As soon as Epoch AI registers the model in its catalog, the daily cron places it in the recently launched band of the page, with a dash where the score goes, because no score exists yet. Once Epoch publishes the ECI, it leaves that band and enters the ranking at whatever position the number gives it. That page measures no models of its own: it aggregates the Epoch AI ECI and the Arena ELO, and that is all it shows.

What are Sonnet 5.5 and Haiku 5.5?

They are the smaller models of the same generation, promised by Anthropic for the coming weeks. They tend to be the ones most used in production, because they handle most tasks at a fraction of the top-tier price. If the price cut seen in Opus 5.5 repeats in them, the impact on anyone running volume will be far larger than today.

Sources
Rodrigo Fávaro

Rodrigo Fávaro

Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.

X @rodmf LinkedIn rodrigofavaro

Want to apply this in your company?

ROO3 diagnoses what can be automated first in your business. The first conversation is free.