Sonnet 5.5 is out: same price as GPT-6 Sol, and at max effort it costs more than Opus
Anthropic kept US$ 2 and US$ 10 per million tokens, the same sheet OpenAI charges for GPT-6 Sol. With the price tied, the effort level decides the bill, and Anthropic's own charts show where the cheap model stops being cheap.
Anthropic shipped Claude Sonnet 5.5 on 28 September 2026 at the same price as Sonnet 5: US$ 2 per million input tokens, US$ 10 per million output and US$ 0.20 for cache reads, exactly the sheet OpenAI charges for GPT-6 Sol. In the figures Anthropic published, it beats Sonnet 5 on all nine tests, passes Opus 5.5 on terminal work (70.6% against 66.4%) and on automation, and trails it on the other seven. What the headline does not show is in the cost charts: in three of the four tests with published cost, there is an Opus 5.5 effort level that scores higher than Sonnet 5.5 at its best and costs less per task. The model pays off at low, medium and high effort, not at max.
What you get from this article
- The identifier claude-sonnet-5-5 has been in Anthropic official catalog since 28 September 2026.
- Same price as Sonnet 5 and, to the cent, as GPT-6 Sol: US$ 2 input, US$ 10 output, US$ 0.20 for cache reads.
- Against Sonnet 5, it wins all nine published tests. On Terminal-Bench 4.0 it goes from 10.3% to 70.6%.
- Against Opus 5.5, it wins two tests (terminal and automation) and loses seven, most of them narrowly.
- At max effort it burns so many tokens that, in three of four tests, an Opus 5.5 level scores higher for less money.
- Anyone using forced tool_choice must change their code before switching models: on Sonnet 5.5 it returns a 400 error.
- Independent measurement is still partial: 75.5% on APEX-Agents, the best result on that test, and 77.8 on LiveBench, below GPT-6 Sol.
What shipped, and where you can already use it
Claude Sonnet 5.5 shipped on 28 September 2026, six days after Opus 5.5 and GPT-6 Sol, which OpenAI launched on that same 22 September. It is the second model in Anthropic's 5.5 family, and Haiku 5.5, the smallest and cheapest, was promised for the coming weeks, with no date.
Confirmation works the usual way, and it is not a screenshot of a table: the identifier claude-sonnet-5-5 is in the official documentation, with published pricing, and the model already runs on Anthropic's API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. The same day it reached GitHub Copilot on the paid plans (Pro, Pro+, Max, Business and Enterprise), billed at list price. It does not show up on the free Copilot plan.
The spec sheet matches its bigger sibling: a 1 million token window, up to 128 thousand output tokens and reliable knowledge through June 2026. Anthropic presents it as the fast, cheap complement to Opus 5.5, strong on well-scoped tasks, bug fixing and documents, slides and spreadsheets. Opus remains, in the company's words, the model for open-ended work that needs sustained judgment.
A curious detail from the announcement: it is the first Sonnet to finish Pokémon Red playing only from screenshots. It sounds like a stunt, but it is an honest long-task test, where the model has to remember what it already did and decide the next step for hours.
Anthropic's table, with the nine tests
Two readings before the numbers. The first is against Sonnet 5, and there is no contest: the new model wins all nine tests, some with jumps that look like typos. On Terminal-Bench 4.0, which measures long command-line work, it went from 10.3% to 70.6%. On Zapier's AutomationBench, from 10.7% to 44.7%. On chart reading, from 15.6% to 61.6%.
The second is against Opus 5.5, which costs twice as much. Sonnet 5.5 wins two tests, terminal and automation, and loses seven, most of them narrowly: two points on GDPval-AA (1,844 against 1,846) and 1.7 points on OSWorld 2.1. The widest gap is SWE-Bench Pro, 81.3% against 89.9%. Even the terminal win deserves care: the system card itself reports a standard error of 2.5 points for Sonnet 5.5 and 2.6 for Opus 5.5 on that test, so the 4.2-point gap is a trend, not a cushion.
Against GPT-6 Sol, Anthropic only had figures on four rows, and Sonnet 5.5 leads on all four. One caveat Anthropic itself notes: OpenAI recently fixed a bug that degraded GPT-6 Sol's image understanding, and some of its scores may not reflect the fix.
| Sonnet 5.5 new | Sonnet 5 | Opus 5.5 | GPT-6 Sol OpenAI | |
|---|---|---|---|---|
| Input price US$ per million tokens | US$ 2 | US$ 2 | US$ 4 | US$ 2 |
| Output price US$ per million tokens | US$ 10 | US$ 10 | US$ 20 | US$ 10 |
| Cache read US$ per million tokens | US$ 0.20 | US$ 0.20 | US$ 0.20 | US$ 0.20 |
| Long terminal work Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | not published |
| Software engineering FrontierCode 1.1 | 52.1% | 42.4% | 54.4% | 49.3% |
| Software engineering CursorBench 4.0 | 55.5% | 34.1% | 57.8% | not published |
| Software engineering SWE-Bench Pro | 81.3% | 63.2% | 89.9% | not published |
| Knowledge work GDPval-AA v2.1 | 1,844 | 1,449 | 1,846 | 1,487 |
| Automation AutomationBench | 44.7% | 10.7% | 42.5% | 32.0% |
| Computer use OSWorld 2.1 | 80.1% | 57.0% | 81.8% | not published |
| Hard knowledge Humanity's Last Exam, with tools | 64.5% | 54.9% | 67.7% | not published |
| Chart reading Chartography | 61.6% | 15.6% | 64.4% | 53.6% |
Same price as GPT-6 Sol: the fight moved elsewhere
Sonnet 5.5 costs the same as Sonnet 5: US$ 2 per million input tokens, US$ 10 output, US$ 0.20 for cache reads and US$ 2.50 for five-minute cache writes. In batch, half. Worth remembering: Sonnet 5 launched at that price as a promotion, with an increase scheduled to US$ 3 and US$ 15 on 1 September. The increase never happened: Anthropic's price sheet says the launch price became the standard price.
OpenAI's pricing page lists exactly the same four numbers for GPT-6 Sol, to the cent. When a token costs the same, the comparison moves and becomes two questions: how many tokens each model burns to deliver the same work, and where each company charges extra.
The second question has an objective answer. OpenAI doubles the input and cache price and charges 1.5 times the output, on the whole request, once the prompt goes past 272 thousand tokens. Anthropic charges the same sheet up to 1 million. For anyone putting a full contract, a manual or a code repository in the prompt, that line weighs more than the headline.
There is a third subtlety: a token is not the same unit at both companies. Each one cuts text into pieces its own way, and Anthropic itself states that the tokenizer it has used since Opus 4.7 produces about 30% more tokens for the same text than the previous one (Sonnet 5.5 kept Sonnet 5's). Equal price per token does not guarantee equal price for the same text. If the subject is new to you, the article on what a token and a context window are walks through the math.
On the first question, Anthropic's charts give a hint. On FrontierCode, measured by Cognition, Sonnet 5.5 at high effort scores 49.4% for US$ 0.42 per task, and GPT-6 Sol at its best scores 49.3% at max, for US$ 2.07. Same score for a fifth of the cost. But the chart is Anthropic's, and the broadest independent test published so far, LiveBench, puts GPT-6 Sol ahead (more on that below).
When two companies charge the same per token, price stops being an argument. What is left is how much each model spends to deliver the same work.
The trap: at max effort, the cheap one gets expensive
The new models have an effort dial with five levels: low, medium, high, xhigh and max. More effort means more reasoning before answering, more tokens and more time. In the Claude app and in Claude Code the default is medium. On the API it is high.
Along with the announcement, Anthropic published the score and cost per task at each level for Sonnet 5.5, Opus 5.5, Sonnet 5 and an OpenAI model, on four tests. Read point by point, those charts show something the announcement text does not say this plainly: at max effort Sonnet 5.5 burns so many tokens that, on three of the four tests, there is an Opus 5.5 level with a higher score and a lower cost per task.
| Sonnet 5.5 medium | Sonnet 5.5 best level | Opus 5.5 that ties or beats it | |
|---|---|---|---|
| Long terminal work Terminal-Bench 4.0, cost per attempt | 28.8% · US$ 0.83 | 70.6% · US$ 12.54 (max) | 66.4% · US$ 7.35 (xhigh, its best) |
| Software engineering FrontierCode 1.1 | 36.5% · US$ 0.24 | 52.1% · US$ 1.59 (xhigh) | 54.6% · US$ 0.80 (medium) |
| Software engineering CursorBench 4.0 | 39.2% · US$ 0.70 | 55.5% · US$ 9.67 (max) | 56.0% · US$ 3.97 (high) |
| Long knowledge-work projects AA-Briefcase v1.1, Elo score | 1,461 · US$ 1.64 | 1,811 · US$ 29.19 (max) | 1,822 · US$ 21.05 (max) |
The clearest case is CursorBench, which uses tasks from real sessions in the Cursor editor. Sonnet 5.5's best is 55.5% for US$ 9.67 per task. Opus 5.5 at high effort scores 56.0% for US$ 3.97, less than half. On AA-Briefcase, Sonnet at max reaches 1,811 points spending US$ 29.19 per task, and Opus at max scores 1,822 for US$ 21.05. On FrontierCode, Opus at medium beats Sonnet's best for half the price.
Comparing level with level, Sonnet 5.5 is the cheaper of the two from low to xhigh, on all four tests. Max is where the order flips: there it costs more per task than Opus 5.5 at max on three of the four, and on FrontierCode the gap is US$ 20.78 against US$ 6.19.
The exception is Terminal-Bench, where Sonnet 5.5 at max scores 70.6%, above every Opus level, paying US$ 12.54 per attempt against US$ 7.35 for the best Opus. And the first column deserves a careful look: at medium, the app default, Sonnet 5.5 scores 28.8% on that test, against 57.6% for Opus 5.5 at the same level.
Anthropic says it its own way: Sonnet complements Opus best at lower settings, and at higher ones delivers comparable results at similar cost. The same points allow a harsher reading, because on three of four tests the cost is not similar, it is higher. And time is a cost too. On the system card's test of questions from health professionals, answers took 12 to 26 seconds from low to xhigh and about 160 seconds at max, to gain between 1 and 1.6 points, within the noise.
The customer testimonials Anthropic chose point the same way, toward volume use. Balyasny, an investment manager, reported about 121 thousand tokens per answer across 2,441 finance tasks, against 497 thousand for Sonnet 5. Base44 counted 3.6 rounds per app built, against 7.7 for Opus 5, across 118 apps. These are numbers from companies invited to test early, but they show where the gain is: doing the same in fewer steps.
Price per token is what is on the sheet. Price per task is what lands on the invoice, and it depends more on the effort dial than on the model name.
What has been measured outside Anthropic
Everything so far was published by the seller. Some tests were run by third parties, such as Cursor, Cognition and Zapier, but Anthropic chose which ones to show. Our AI Benchmark gathers measurements that do not depend on the maker, and Sonnet 5.5 entered it on its own the morning after launch, on the first run after Epoch AI registered it. As of this writing, this is what exists:
| Sonnet 5.5 new | Opus 5.5 | GPT-6 Sol OpenAI | Sonnet 5 | |
|---|---|---|---|---|
| Overall average, 23 tasks LiveBench, 0 to 100 | 77.8 | 83.2 | 79.3 | 76.0 |
| Agentic office work APEX-Agents | 75.5% | 73.5% | 54.3% | 54.5% |
| Scientific coding SciCode | 61.0% | 66.9% | 57.6% | 53.6% |
| Research-level physics CritPt | 31.4% | 31.7% | 30.9% | 16.9% |
The strongest data point in the model's favor is APEX-Agents, which measures office work carried out by an agent: 75.5%, the best result ever recorded on the test, above Opus 5.5 (73.5%) and well above GPT-6 Sol (54.3%). It matches the story the maker tells: the jump is in multi-step work with tools.
LiveBench, the broad test, tells a different story: 77.8, only 1.7 points above Sonnet 5 and below GPT-6 Sol. And the number looks unstable. At max effort Sonnet 5.5 scored 75.7, less than at xhigh and less than Sonnet 5, and the agentic coding category swings from 56.3 at max to 39.3 at xhigh, when Sonnet 5 had 59.4. That fits no other measurement of the model. Read it as provisional until LiveBench's next update.
The two measurements the AI Benchmark uses to rank a model are still missing: Epoch AI's composite index (ECI) and the Arena human-preference score. That is normal: both usually arrive days after a launch. When they do, the model moves into the ranking on its own, with nobody typing a number.
What breaks for code already running on Sonnet 5
Changing the model name in your config is not enough. The documentation lists five changes that break code written for Sonnet 5, and two of them hit common integrations, such as automated customer service and data extraction.
The most serious is the end of forced tool use. Many systems set tool_choice to any or tool to force the model to answer in a fixed format, such as the customer's name, phone and intent. On Sonnet 5.5 that request comes back with a 400 error. The official replacement is to leave it on auto and turn on strict tool use, or to use structured outputs, and to say in the prompt when the tool applies.
- Forced tool use:
tool_choiceset toanyortoolreturns a 400 error. Useautowith strict mode. - Thinking turned off:
disabledno longer exists and errors out. The lowest setting is nowbetween_tools, accepted only up to high effort. - Edited history: reasoning is bound to the model and to the conversation. On accounts created on or after 31 August 2026, replaying a reasoning block after changing the system prompt, the tools or an earlier message returns a 400 error. The conversation must only grow, never be rewritten.
- Computer use: on Anthropic's API and Google Cloud, the older
computer_20251124tool is not accepted. You have to move to the new toolset. - Advisor model: the advisor tool no longer accepts Opus 4.8, Opus 4.7 or Sonnet 5 in that role.
And there is one change that throws no error at all, which makes it the trickiest: the text the model writes between one tool call and the next now comes back inside reasoning blocks, empty by default. A screen that used to show "checking the stock" to the customer goes silent, with no alarm. The fix is one configuration line, but only someone who knows about the problem will look for it.
The documentation also warns that effort levels were recalibrated: Sonnet 5.5's medium does not think as much as Sonnet 5's medium. Carrying the old setting straight over to the new model is the fastest way to pay a different amount without noticing.
Safety, and what the system card itself admits
It is the first Sonnet launched with the cybersecurity safeguards of the company's strongest models, because its capability in that area came close to Opus 5's. In practice, fixing bugs day to day works as usual, but high-risk security requests visibly fall back to Sonnet 5. It is also the first Sonnet with blocking against anyone trying to extract the model's reasoning to train a competitor.
The system card runs 148 pages and includes less flattering findings worth repeating. On a closed-book knowledge test, Sonnet 5.5 gets more answers right than Sonnet 5, but is slightly more likely to state a wrong answer instead of saying it does not know, and on net it only beats Sonnet 5 among the Claude models compared. Its reasoning text is less legible than recent models'. And its warmth and humor came out slightly below Sonnet 5's.
On the good side, in the automated behavioral audit, with about 1,850 scenarios, it matches or improves on Sonnet 5 on most measures, has the lowest rate of attempted sandbox escape among the models tested and is the most robust Sonnet yet against malicious instructions hidden in pages and files, especially in coding and browser use. The usual caveat: the seller did the measuring. For an agent that touches real systems, tight permissions are still worth more than good behavior.
What to do this week
If you use AI through the app, on a monthly plan, the news is good and asks nothing of you: the middle model got faster and much more capable. You do not pick a model by token price.
If you built something that calls the API, the math changed shape, and the decisions fit in five lines.
- Before switching models, search your code for forced
tool_choiceand disabledthinking. Both become 400 errors on Sonnet 5.5. - Do not copy Sonnet 5's effort level. Rerun the test: the documentation suggests starting at high in general, at medium for agents with well-specified tasks and at low or medium for chat.
- Before raising Sonnet 5.5 to xhigh or max, run the same tasks on Opus 5.5 at medium or high. By Anthropic's own charts, it may come out better and cheaper.
- If you were choosing between Sonnet and GPT-6 Sol on price, the price is now tied. Compare cost per task on your own tasks, and check whether your prompts go past 272 thousand tokens, where OpenAI charges more.
- If the task is simple and high volume, such as classifying, summarizing and extracting, measure at low. And consider waiting for Haiku 5.5 before migrating twice.
None of these decisions gets settled by reading benchmarks, this text included. They get settled by running ten of your real tasks on the candidate models and looking at two columns: accuracy and cost. It is an afternoon of work, and it is worth more than any launch table.
What cannot be said yet
First, where Sonnet 5.5 stands among models. Without Epoch AI's ECI and without the Arena score, what exists independently is a handful of individual tests and a LiveBench score that looks provisional. That is too little to say whether it beats GPT-6 Sol for general use.
Second, what changes in the Claude app. As of this writing, Anthropic has not detailed which plans now use Sonnet 5.5 by default. Through the API and the clouds, it has been available to everyone since day one.
Third, Haiku 5.5. It sets the price of the cheap tier, the one that runs triage, summaries and customer service at volume. Last week, Opus 5.5 brought down the price at the top. This week, Sonnet 5.5 held the middle price and raised what it delivers. If Haiku repeats the move, the bill for anyone running volume changes more than it has so far.
Frequently asked questions
Is Sonnet 5.5 better than Opus 5.5?
Overall, no. In Anthropic's table, Sonnet 5.5 wins two of the nine tests (terminal work and automation) and loses seven, most narrowly, and the company itself says Opus 5.5 remains stronger for open-ended, long work. On one independent measurement, APEX-Agents, Sonnet 5.5 scored 75.5% against 73.5% for Opus 5.5. It costs half per token: US$ 2 and US$ 10, against US$ 4 and US$ 20.
How much does Sonnet 5.5 cost?
US$ 2 per million input tokens and US$ 10 per million output. Cache reads cost US$ 0.20, cache writes US$ 2.50 (five minutes) or US$ 4 (one hour), and batch processing is half price: US$ 1 and US$ 5. It is the same price as Sonnet 5, and it applies across the full 1 million token window.
Sonnet 5.5 or GPT-6 Sol: which one should I pick?
The price per token is identical: US$ 2 input, US$ 10 output and US$ 0.20 cache on both. In the tests Anthropic published, Sonnet 5.5 leads on the four where both have figures. On LiveBench, which is independent, GPT-6 Sol leads, 79.3 against 77.8. Above 272 thousand tokens per request, OpenAI charges more and Anthropic does not. The right choice comes from running your own tasks on both and comparing accuracy and cost per task.
Why can Sonnet 5.5 at max effort cost more than Opus 5.5?
Because at max it uses far more tokens. On CursorBench, Sonnet 5.5 spent on average 271,920 tokens per task at max, against 37,391 at high, more than seven times as many, to go from 47.8% to 55.5%. As a result, Opus 5.5 at high effort scores 56.0% for US$ 3.97 per task, while Sonnet 5.5 at its best costs US$ 9.67. Comparing level with level, Sonnet is the cheaper of the two from low to xhigh; it is at max that it starts costing more.
Do I need to change my code to use Sonnet 5.5?
If your code uses forced tool use (tool_choice any or tool), turns thinking off with disabled, edits the conversation history before resending it, uses the older computer use tool or has Sonnet 5 as an advisor, yes: those cases return a 400 error. Otherwise, switch the identifier to claude-sonnet-5-5 and rerun your effort-level test, because the levels were recalibrated.
Is Sonnet 5.5 already on the ROO3 AI Benchmark?
Yes, since the morning after launch, when Epoch AI registered it, and with no figure from the maker. It shows up as a recent release, with the LiveBench average and the individual tests Epoch AI has already published, each next to the best result on that test. It has no position in the ranking yet, which is ordered by Epoch AI's composite index. When that index comes out, it enters on its own at whatever position the number gives it.
When is Haiku 5.5 coming?
Anthropic only said that Haiku 5.5, built for high-volume, cost-sensitive use, joins the 5.5 family in the coming weeks. No date or price has been published. Until then, the current Haiku is 4.5, at US$ 1 and US$ 5 per million tokens.
- Anthropic, Claude Sonnet 5.5 announcement (benchmark table, score-versus-cost charts and pricing)
- Anthropic, Claude Sonnet 5.5 System Card (148-page technical report)
- Claude Platform Docs, What's new in Claude Sonnet 5.5 (breaking changes)
- Claude Platform Docs, Pricing (Anthropic official price sheet)
- OpenAI, API pricing (GPT-6 Sol pricing, including long context)
- GitHub Changelog, Claude Sonnet 5.5 in GitHub Copilot
- LiveBench, leaderboard and CC BY-SA 4.0 license
- Epoch AI, benchmark data (APEX-Agents, SciCode and CritPt, under CC BY 4.0)
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

Opus 5.5: the expensive one got better and cheaper at the same time
The official Claude Opus 5.5 benchmarks against Fable 5.1, Opus 5 and GPT-6 Astra, the price per token, the 60% drop in...
11 min read
Grok 4.7: the cheapest model at the table wins two tests out of seven
The official Grok 4.7 benchmarks against GPT-5.6 Sol and Claude Fable 5.1, the price per token, the fine print that...
8 min read
Which AI is best today: how to compare without following opinion
How to compare AI models using public measurement instead of opinion: what the Epoch AI ECI is, what the Arena ELO is...
9 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.