Tools

Sonnet 5.5 is out: same price as GPT-6 Sol, and at max effort it costs more than Opus

Anthropic kept US$ 2 and US$ 10 per million tokens, the same sheet OpenAI charges for GPT-6 Sol. With the price tied, the effort level decides the bill, and Anthropic's own charts show where the cheap model stops being cheap.

Rodrigo Fávaro, fundador da ROO3
Rodrigo Fávaro Founder of ROO3
·15 min read
Comparison table showing Claude Sonnet 5.5 against Sonnet 5, Opus 5.5 and GPT-6 Sol across seven performance tests, with the price per million tokens.
Short answer

Anthropic shipped Claude Sonnet 5.5 on 28 September 2026 at the same price as Sonnet 5: US$ 2 per million input tokens, US$ 10 per million output and US$ 0.20 for cache reads, exactly the sheet OpenAI charges for GPT-6 Sol. In the figures Anthropic published, it beats Sonnet 5 on all nine tests, passes Opus 5.5 on terminal work (70.6% against 66.4%) and on automation, and trails it on the other seven. What the headline does not show is in the cost charts: in three of the four tests with published cost, there is an Opus 5.5 effort level that scores higher than Sonnet 5.5 at its best and costs less per task. The model pays off at low, medium and high effort, not at max.

What you get from this article

  • The identifier claude-sonnet-5-5 has been in Anthropic official catalog since 28 September 2026.
  • Same price as Sonnet 5 and, to the cent, as GPT-6 Sol: US$ 2 input, US$ 10 output, US$ 0.20 for cache reads.
  • Against Sonnet 5, it wins all nine published tests. On Terminal-Bench 4.0 it goes from 10.3% to 70.6%.
  • Against Opus 5.5, it wins two tests (terminal and automation) and loses seven, most of them narrowly.
  • At max effort it burns so many tokens that, in three of four tests, an Opus 5.5 level scores higher for less money.
  • Anyone using forced tool_choice must change their code before switching models: on Sonnet 5.5 it returns a 400 error.
  • Independent measurement is still partial: 75.5% on APEX-Agents, the best result on that test, and 77.8 on LiveBench, below GPT-6 Sol.

What shipped, and where you can already use it

Claude Sonnet 5.5 shipped on 28 September 2026, six days after Opus 5.5 and GPT-6 Sol, which OpenAI launched on that same 22 September. It is the second model in Anthropic's 5.5 family, and Haiku 5.5, the smallest and cheapest, was promised for the coming weeks, with no date.

Confirmation works the usual way, and it is not a screenshot of a table: the identifier claude-sonnet-5-5 is in the official documentation, with published pricing, and the model already runs on Anthropic's API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. The same day it reached GitHub Copilot on the paid plans (Pro, Pro+, Max, Business and Enterprise), billed at list price. It does not show up on the free Copilot plan.

The spec sheet matches its bigger sibling: a 1 million token window, up to 128 thousand output tokens and reliable knowledge through June 2026. Anthropic presents it as the fast, cheap complement to Opus 5.5, strong on well-scoped tasks, bug fixing and documents, slides and spreadsheets. Opus remains, in the company's words, the model for open-ended work that needs sustained judgment.

A curious detail from the announcement: it is the first Sonnet to finish Pokémon Red playing only from screenshots. It sounds like a stunt, but it is an honest long-task test, where the model has to remember what it already did and decide the next step for hours.

Anthropic's table, with the nine tests

Two readings before the numbers. The first is against Sonnet 5, and there is no contest: the new model wins all nine tests, some with jumps that look like typos. On Terminal-Bench 4.0, which measures long command-line work, it went from 10.3% to 70.6%. On Zapier's AutomationBench, from 10.7% to 44.7%. On chart reading, from 15.6% to 61.6%.

The second is against Opus 5.5, which costs twice as much. Sonnet 5.5 wins two tests, terminal and automation, and loses seven, most of them narrowly: two points on GDPval-AA (1,844 against 1,846) and 1.7 points on OSWorld 2.1. The widest gap is SWE-Bench Pro, 81.3% against 89.9%. Even the terminal win deserves care: the system card itself reports a standard error of 2.5 points for Sonnet 5.5 and 2.6 for Opus 5.5 on that test, so the 4.2-point gap is a trend, not a cushion.

Against GPT-6 Sol, Anthropic only had figures on four rows, and Sonnet 5.5 leads on all four. One caveat Anthropic itself notes: OpenAI recently fixed a bug that degraded GPT-6 Sol's image understanding, and some of its scores may not reflect the fix.

Sonnet 5.5 against Sonnet 5, Opus 5.5 and GPT-6 Sol, in the figures Anthropic published
Sonnet 5.5 new Sonnet 5 Opus 5.5 GPT-6 Sol OpenAI
Input price US$ per million tokens US$ 2 US$ 2 US$ 4 US$ 2
Output price US$ per million tokens US$ 10 US$ 10 US$ 20 US$ 10
Cache read US$ per million tokens US$ 0.20 US$ 0.20 US$ 0.20 US$ 0.20
Long terminal work Terminal-Bench 4.0 70.6% 10.3% 66.4% not published
Software engineering FrontierCode 1.1 52.1% 42.4% 54.4% 49.3%
Software engineering CursorBench 4.0 55.5% 34.1% 57.8% not published
Software engineering SWE-Bench Pro 81.3% 63.2% 89.9% not published
Knowledge work GDPval-AA v2.1 1,844 1,449 1,846 1,487
Automation AutomationBench 44.7% 10.7% 42.5% 32.0%
Computer use OSWorld 2.1 80.1% 57.0% 81.8% not published
Hard knowledge Humanity's Last Exam, with tools 64.5% 54.9% 67.7% not published
Chart reading Chartography 61.6% 15.6% 64.4% 53.6%
Figures from the Sonnet 5.5 announcement and system card. On FrontierCode, Anthropic published two values for Sonnet 5.5: 46.2% at max effort and 52.1% at xhigh, and explained that at max the model more often made out-of-scope changes, which the test penalizes. We show the higher one, as we do on our AI Benchmark. AutomationBench comes from the system card, with Opus 5.5 re-run with the fallback model enabled (42.5%; it was 40.0% at its own launch). GPT-6 Sol pricing is from OpenAI's pricing page.

Same price as GPT-6 Sol: the fight moved elsewhere

Sonnet 5.5 costs the same as Sonnet 5: US$ 2 per million input tokens, US$ 10 output, US$ 0.20 for cache reads and US$ 2.50 for five-minute cache writes. In batch, half. Worth remembering: Sonnet 5 launched at that price as a promotion, with an increase scheduled to US$ 3 and US$ 15 on 1 September. The increase never happened: Anthropic's price sheet says the launch price became the standard price.

OpenAI's pricing page lists exactly the same four numbers for GPT-6 Sol, to the cent. When a token costs the same, the comparison moves and becomes two questions: how many tokens each model burns to deliver the same work, and where each company charges extra.

The second question has an objective answer. OpenAI doubles the input and cache price and charges 1.5 times the output, on the whole request, once the prompt goes past 272 thousand tokens. Anthropic charges the same sheet up to 1 million. For anyone putting a full contract, a manual or a code repository in the prompt, that line weighs more than the headline.

There is a third subtlety: a token is not the same unit at both companies. Each one cuts text into pieces its own way, and Anthropic itself states that the tokenizer it has used since Opus 4.7 produces about 30% more tokens for the same text than the previous one (Sonnet 5.5 kept Sonnet 5's). Equal price per token does not guarantee equal price for the same text. If the subject is new to you, the article on what a token and a context window are walks through the math.

On the first question, Anthropic's charts give a hint. On FrontierCode, measured by Cognition, Sonnet 5.5 at high effort scores 49.4% for US$ 0.42 per task, and GPT-6 Sol at its best scores 49.3% at max, for US$ 2.07. Same score for a fifth of the cost. But the chart is Anthropic's, and the broadest independent test published so far, LiveBench, puts GPT-6 Sol ahead (more on that below).

When two companies charge the same per token, price stops being an argument. What is left is how much each model spends to deliver the same work.

The trap: at max effort, the cheap one gets expensive

The new models have an effort dial with five levels: low, medium, high, xhigh and max. More effort means more reasoning before answering, more tokens and more time. In the Claude app and in Claude Code the default is medium. On the API it is high.

Along with the announcement, Anthropic published the score and cost per task at each level for Sonnet 5.5, Opus 5.5, Sonnet 5 and an OpenAI model, on four tests. Read point by point, those charts show something the announcement text does not say this plainly: at max effort Sonnet 5.5 burns so many tokens that, on three of the four tests, there is an Opus 5.5 level with a higher score and a lower cost per task.

Score and cost per task on four tests
Sonnet 5.5 medium Sonnet 5.5 best level Opus 5.5 that ties or beats it
Long terminal work Terminal-Bench 4.0, cost per attempt 28.8% · US$ 0.83 70.6% · US$ 12.54 (max) 66.4% · US$ 7.35 (xhigh, its best)
Software engineering FrontierCode 1.1 36.5% · US$ 0.24 52.1% · US$ 1.59 (xhigh) 54.6% · US$ 0.80 (medium)
Software engineering CursorBench 4.0 39.2% · US$ 0.70 55.5% · US$ 9.67 (max) 56.0% · US$ 3.97 (high)
Long knowledge-work projects AA-Briefcase v1.1, Elo score 1,461 · US$ 1.64 1,811 · US$ 29.19 (max) 1,822 · US$ 21.05 (max)
Points from the score-versus-cost charts in Anthropic's announcement, in dollars per task at list price. The dot marks the highest score in the row. On Terminal-Bench, no Opus 5.5 level reaches Sonnet 5.5's 70.6%, so the column shows Opus's best result.

The clearest case is CursorBench, which uses tasks from real sessions in the Cursor editor. Sonnet 5.5's best is 55.5% for US$ 9.67 per task. Opus 5.5 at high effort scores 56.0% for US$ 3.97, less than half. On AA-Briefcase, Sonnet at max reaches 1,811 points spending US$ 29.19 per task, and Opus at max scores 1,822 for US$ 21.05. On FrontierCode, Opus at medium beats Sonnet's best for half the price.

Comparing level with level, Sonnet 5.5 is the cheaper of the two from low to xhigh, on all four tests. Max is where the order flips: there it costs more per task than Opus 5.5 at max on three of the four, and on FrontierCode the gap is US$ 20.78 against US$ 6.19.

The exception is Terminal-Bench, where Sonnet 5.5 at max scores 70.6%, above every Opus level, paying US$ 12.54 per attempt against US$ 7.35 for the best Opus. And the first column deserves a careful look: at medium, the app default, Sonnet 5.5 scores 28.8% on that test, against 57.6% for Opus 5.5 at the same level.

Anthropic says it its own way: Sonnet complements Opus best at lower settings, and at higher ones delivers comparable results at similar cost. The same points allow a harsher reading, because on three of four tests the cost is not similar, it is higher. And time is a cost too. On the system card's test of questions from health professionals, answers took 12 to 26 seconds from low to xhigh and about 160 seconds at max, to gain between 1 and 1.6 points, within the noise.

The customer testimonials Anthropic chose point the same way, toward volume use. Balyasny, an investment manager, reported about 121 thousand tokens per answer across 2,441 finance tasks, against 497 thousand for Sonnet 5. Base44 counted 3.6 rounds per app built, against 7.7 for Opus 5, across 118 apps. These are numbers from companies invited to test early, but they show where the gain is: doing the same in fewer steps.

Price per token is what is on the sheet. Price per task is what lands on the invoice, and it depends more on the effort dial than on the model name.

What has been measured outside Anthropic

Everything so far was published by the seller. Some tests were run by third parties, such as Cursor, Cognition and Zapier, but Anthropic chose which ones to show. Our AI Benchmark gathers measurements that do not depend on the maker, and Sonnet 5.5 entered it on its own the morning after launch, on the first run after Epoch AI registered it. As of this writing, this is what exists:

Sonnet 5.5 in measurements published outside the announcement, as of 29 September 2026
Sonnet 5.5 new Opus 5.5 GPT-6 Sol OpenAI Sonnet 5
Overall average, 23 tasks LiveBench, 0 to 100 77.8 83.2 79.3 76.0
Agentic office work APEX-Agents 75.5% 73.5% 54.3% 54.5%
Scientific coding SciCode 61.0% 66.9% 57.6% 53.6%
Research-level physics CritPt 31.4% 31.7% 30.9% 16.9%
LiveBench, under CC BY-SA 4.0, with the average recomputed using the same formula as the site. The other three tests come from the Epoch AI package, under CC BY 4.0. In every row, each model's best effort level counts. ROO3 does not measure any model.

The strongest data point in the model's favor is APEX-Agents, which measures office work carried out by an agent: 75.5%, the best result ever recorded on the test, above Opus 5.5 (73.5%) and well above GPT-6 Sol (54.3%). It matches the story the maker tells: the jump is in multi-step work with tools.

LiveBench, the broad test, tells a different story: 77.8, only 1.7 points above Sonnet 5 and below GPT-6 Sol. And the number looks unstable. At max effort Sonnet 5.5 scored 75.7, less than at xhigh and less than Sonnet 5, and the agentic coding category swings from 56.3 at max to 39.3 at xhigh, when Sonnet 5 had 59.4. That fits no other measurement of the model. Read it as provisional until LiveBench's next update.

The two measurements the AI Benchmark uses to rank a model are still missing: Epoch AI's composite index (ECI) and the Arena human-preference score. That is normal: both usually arrive days after a launch. When they do, the model moves into the ranking on its own, with nobody typing a number.

What breaks for code already running on Sonnet 5

Changing the model name in your config is not enough. The documentation lists five changes that break code written for Sonnet 5, and two of them hit common integrations, such as automated customer service and data extraction.

The most serious is the end of forced tool use. Many systems set tool_choice to any or tool to force the model to answer in a fixed format, such as the customer's name, phone and intent. On Sonnet 5.5 that request comes back with a 400 error. The official replacement is to leave it on auto and turn on strict tool use, or to use structured outputs, and to say in the prompt when the tool applies.

And there is one change that throws no error at all, which makes it the trickiest: the text the model writes between one tool call and the next now comes back inside reasoning blocks, empty by default. A screen that used to show "checking the stock" to the customer goes silent, with no alarm. The fix is one configuration line, but only someone who knows about the problem will look for it.

The documentation also warns that effort levels were recalibrated: Sonnet 5.5's medium does not think as much as Sonnet 5's medium. Carrying the old setting straight over to the new model is the fastest way to pay a different amount without noticing.

Safety, and what the system card itself admits

It is the first Sonnet launched with the cybersecurity safeguards of the company's strongest models, because its capability in that area came close to Opus 5's. In practice, fixing bugs day to day works as usual, but high-risk security requests visibly fall back to Sonnet 5. It is also the first Sonnet with blocking against anyone trying to extract the model's reasoning to train a competitor.

The system card runs 148 pages and includes less flattering findings worth repeating. On a closed-book knowledge test, Sonnet 5.5 gets more answers right than Sonnet 5, but is slightly more likely to state a wrong answer instead of saying it does not know, and on net it only beats Sonnet 5 among the Claude models compared. Its reasoning text is less legible than recent models'. And its warmth and humor came out slightly below Sonnet 5's.

On the good side, in the automated behavioral audit, with about 1,850 scenarios, it matches or improves on Sonnet 5 on most measures, has the lowest rate of attempted sandbox escape among the models tested and is the most robust Sonnet yet against malicious instructions hidden in pages and files, especially in coding and browser use. The usual caveat: the seller did the measuring. For an agent that touches real systems, tight permissions are still worth more than good behavior.

What to do this week

If you use AI through the app, on a monthly plan, the news is good and asks nothing of you: the middle model got faster and much more capable. You do not pick a model by token price.

If you built something that calls the API, the math changed shape, and the decisions fit in five lines.

None of these decisions gets settled by reading benchmarks, this text included. They get settled by running ten of your real tasks on the candidate models and looking at two columns: accuracy and cost. It is an afternoon of work, and it is worth more than any launch table.

What cannot be said yet

First, where Sonnet 5.5 stands among models. Without Epoch AI's ECI and without the Arena score, what exists independently is a handful of individual tests and a LiveBench score that looks provisional. That is too little to say whether it beats GPT-6 Sol for general use.

Second, what changes in the Claude app. As of this writing, Anthropic has not detailed which plans now use Sonnet 5.5 by default. Through the API and the clouds, it has been available to everyone since day one.

Third, Haiku 5.5. It sets the price of the cheap tier, the one that runs triage, summaries and customer service at volume. Last week, Opus 5.5 brought down the price at the top. This week, Sonnet 5.5 held the middle price and raised what it delivers. If Haiku repeats the move, the bill for anyone running volume changes more than it has so far.

Frequently asked questions

Is Sonnet 5.5 better than Opus 5.5?

Overall, no. In Anthropic's table, Sonnet 5.5 wins two of the nine tests (terminal work and automation) and loses seven, most narrowly, and the company itself says Opus 5.5 remains stronger for open-ended, long work. On one independent measurement, APEX-Agents, Sonnet 5.5 scored 75.5% against 73.5% for Opus 5.5. It costs half per token: US$ 2 and US$ 10, against US$ 4 and US$ 20.

How much does Sonnet 5.5 cost?

US$ 2 per million input tokens and US$ 10 per million output. Cache reads cost US$ 0.20, cache writes US$ 2.50 (five minutes) or US$ 4 (one hour), and batch processing is half price: US$ 1 and US$ 5. It is the same price as Sonnet 5, and it applies across the full 1 million token window.

Sonnet 5.5 or GPT-6 Sol: which one should I pick?

The price per token is identical: US$ 2 input, US$ 10 output and US$ 0.20 cache on both. In the tests Anthropic published, Sonnet 5.5 leads on the four where both have figures. On LiveBench, which is independent, GPT-6 Sol leads, 79.3 against 77.8. Above 272 thousand tokens per request, OpenAI charges more and Anthropic does not. The right choice comes from running your own tasks on both and comparing accuracy and cost per task.

Why can Sonnet 5.5 at max effort cost more than Opus 5.5?

Because at max it uses far more tokens. On CursorBench, Sonnet 5.5 spent on average 271,920 tokens per task at max, against 37,391 at high, more than seven times as many, to go from 47.8% to 55.5%. As a result, Opus 5.5 at high effort scores 56.0% for US$ 3.97 per task, while Sonnet 5.5 at its best costs US$ 9.67. Comparing level with level, Sonnet is the cheaper of the two from low to xhigh; it is at max that it starts costing more.

Do I need to change my code to use Sonnet 5.5?

If your code uses forced tool use (tool_choice any or tool), turns thinking off with disabled, edits the conversation history before resending it, uses the older computer use tool or has Sonnet 5 as an advisor, yes: those cases return a 400 error. Otherwise, switch the identifier to claude-sonnet-5-5 and rerun your effort-level test, because the levels were recalibrated.

Is Sonnet 5.5 already on the ROO3 AI Benchmark?

Yes, since the morning after launch, when Epoch AI registered it, and with no figure from the maker. It shows up as a recent release, with the LiveBench average and the individual tests Epoch AI has already published, each next to the best result on that test. It has no position in the ranking yet, which is ordered by Epoch AI's composite index. When that index comes out, it enters on its own at whatever position the number gives it.

When is Haiku 5.5 coming?

Anthropic only said that Haiku 5.5, built for high-volume, cost-sensitive use, joins the 5.5 family in the coming weeks. No date or price has been published. Until then, the current Haiku is 4.5, at US$ 1 and US$ 5 per million tokens.

Sources
Rodrigo Fávaro

Rodrigo Fávaro

Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.

X @rodmf LinkedIn rodrigofavaro

Want to apply this in your company?

ROO3 diagnoses what can be automated first in your business. The first conversation is free.