Grok 4.7 is out: it wins two of seven tests, and costs five times less than the model that wins the rest
xAI published the comparison against GPT-5.6 and Claude Fable 5.1. The number that actually decides is not highlighted in the table: it is what happens to the price once your text passes 200,000 tokens.
Grok 4.7 entered xAI catalog in September 2026 at US$ 2 per million input tokens and US$ 6 per million output, the same price as 4.6 and a fraction of what its direct competitors charge. In the comparison xAI published itself, against GPT-5.6 Sol and Claude Fable 5.1, it wins two of the seven tests: legal work and electrical engineering. It loses the four that measure coding and long autonomous work. What the table does not show is that the price doubles once a conversation passes 200,000 tokens, and that one number in its column was measured at a different effort setting from the rest.
What you get from this article
- The model is already in xAI official catalog as grok-4.7, flagged as its most capable.
- Price: US$ 2 per million input tokens and US$ 6 per million output, up to 200,000 tokens.
- Above 200,000 tokens the price DOUBLES, to US$ 4 and US$ 12. That does not appear in the launch table.
- It wins 2 of 7 tests: Harvey, for legal work, and EEBench, for electrical engineering.
- It loses 4 of 7 to Claude Fable 5.1, with the widest gap in terminal work: 38.0% against 57.9%.
- The 71.0% on DeepSWE carries an asterisk: it was measured at "high" effort, not the "xhigh" used elsewhere in that column.
- Musk himself graded the model as "roughly on par with Opus 5.0, not 5.1".
What shipped, and how to confirm it without trusting a screenshot
Grok 4.7 spent weeks being announced before it existed. On 2 September 2026 Elon Musk said it would ship in ten days. The date passed, and he wrote that the model needed "a few more days to cook". By 18 September there was still no release, and no model identifier in xAI documentation.
Now there is. The way to confirm that is not the screenshot going around on X. It is xAI own documentation: the identifier grok-4.7 sits in the model catalog, flagged as the most capable model in the house and recommended for code and chat. That is the test that counts, because anyone can assemble a screenshot of a table.
Alongside the model came the official comparison against GPT-5.6 Sol, from OpenAI, and Claude Fable 5.1, from Anthropic. That is what is reproduced below, with every number checked against the announcement.
The table, with all seven tests
Before reading it: each model was measured at a different effort setting, and that is stated right under each column heading. Effort, in these models, is how much the system thinks before answering, and it moves both the result and the cost.
| Grok 4.7 xhigh | Grok 4.6 high | GPT-5.6 Sol max | Fable 5.1 max | |
|---|---|---|---|---|
| Input price US$ per million tokens | US$ 2 | US$ 2 | US$ 4 | US$ 10 |
| Output price US$ per million tokens | US$ 6 | US$ 6 | US$ 20 | US$ 50 |
| Software engineering CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| Software engineering DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| Multi-hour office work AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Multi-hour terminal work Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Legal work Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| Clinical reasoning HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
| Electrical engineering EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
Where it wins, and what the two tests have in common
Neither of Grok 4.7 two wins is in coding. Both are in domain knowledge.
On Harvey, which measures real legal work, it scores 19.6%. Fable 5.1 scores 6.7% and GPT-5.6 scores 2.5%. That is almost three times the runner-up, and it is not a new advantage: Grok 4.6 already scored 15.8% on the same test, also well ahead of the others.
On EEBench, for electrical engineering, it is 64.0% against 56.4% for Fable and 39.4% for GPT-5.6. Here the internal jump stands out: 4.6 scored 53.0%.
It is worth recording what the absolute numbers say, and not just who won. 19.6% on legal work is low in absolute terms. The test is hard by design, and nobody is close to solving it. Reading "Grok wins at legal" as "you can use Grok instead of a lawyer" is exactly the conclusion the table does not license.
Winning a test where everyone does badly is not the same as being good at it. It tells you who is least far away, not who arrived.
Where it loses, and that is where most people work
Of the seven rows, Claude Fable 5.1 wins four, and they are precisely the ones tied to writing code and carrying out long tasks unattended.
The hardest gap is on Terminal-Bench 4.0, which measures long stretches of work inside a terminal: 38.0% for Grok against 57.9% for Fable. Almost twenty points.
On the other hand, that is also where Grok 4.7 improved most over its predecessor. Version 4.6 scored 20.3% on that test. Going to 38.0% is an 87% relative gain, the largest in the whole table. The model improved a great deal at what used to be its weakest point, and it is still behind.
On CursorBench 4.0 it is 46.3% against 51.8% for Fable. On AA Briefcase, which measures multi-hour office work, 1,657 against 1,678, essentially a tie. On clinical reasoning, 56.7% against 62.1%.
There is one reading detail about CursorBench worth carrying. When Grok 4.5 launched, Cursor itself reported that an old snapshot of its codebase had accidentally made it into the model training, which gave an advantage on that specific test. That was with 4.5, not 4.7, and there is no indication it repeats. But it is a useful reminder that a benchmark is made of data, and data leaks.
The price, and the line the launch table leaves out
Here is the most useful piece of information in this whole article, and it was not in the image that circulated.
The US$ 2 input and US$ 6 output apply up to 200,000 tokens. Past that, xAI documentation shows the price doubling: US$ 4 and US$ 12. Anyone using the model to read a full contract, a large codebase or a long conversation history will land in that bracket without noticing, because billing is per request and nobody counts tokens in their head.
Even doubled, it is still the cheapest at the table. On output, which is where the money actually goes, Fable 5.1 costs US$ 50 per million against US$ 6 for Grok. That is more than eight times. GPT-5.6 Sol costs US$ 20, more than three times.
So the practical question is not "which one is best". It is: are those twenty points on Terminal-Bench worth paying eight times more on output? For writing email, summarizing meetings and answering customers, almost certainly not. For an agent running unattended for hours inside your codebase, almost certainly yes, because there the mistake costs more than the tokens.
List price is about the model. Real price is about the task. A model eight times more expensive that gets it right the first time can come out cheaper than a cheap one that needs three attempts.
How to read a vendor benchmark without fooling yourself
This table was published by xAI, about the xAI model. That does not make it false, and the numbers match what the company itself released. But it changes how you read it, and three things deserve attention.
First: each column sits at a different effort setting. Grok 4.7 appears at "xhigh", 4.6 at "high", the competitors at "max". This is not like for like. It is each model at the setting the publisher chose to show.
Second: there is an asterisk in the middle of the column. The 71.0% on DeepSWE was measured at "high", not the "xhigh" of the other rows. It is an honest note, and it is there. But it means that row is not comparable with its neighbours in the same column.
Third: whoever publishes chooses the tests. Seven tests were shown. There is no way to know how many were run. That holds for every launch table, from every company, and it is the reason to check independent evaluations before deciding anything.
There is a fourth source, and it is the least suspect of all, because it argues against the publisher. On 14 September, before launch, Musk himself graded the model as "roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others". That is exactly what the table shows.
What it changes for people using AI at work
Almost nothing changes this week, and that is the honest answer. What changes is the bill for anyone paying per token at volume.
If you use AI through a chat interface on a monthly plan, none of this reaches you: you do not pick a model by token price. If you built something that calls the API, the table has direct consequences, and they fit into four decisions.
- High volume, low risk work such as classifying messages, summarizing and extracting data: the cheapest one handles it, and the quality gap does not show.
- An agent running unattended for a long time, touching code or systems: pay for the best. The mistake costs more than the tokens.
- If your text regularly passes 200,000 tokens, redo the maths with the doubled price, not the list price.
- Before switching models because of a benchmark, run YOUR task on both. Twenty points on a standardized test can become zero in your case, or the reverse.
Frequently asked questions
Is Grok 4.7 better than Claude and GPT?
It depends on what you are doing. In the comparison xAI published, Grok 4.7 wins two of seven tests, the ones for legal work and electrical engineering. Claude Fable 5.1 wins four, including both coding tests and the one for long terminal work. GPT-5.6 Sol wins one. Grok is by a wide margin the cheapest of the three.
How much does Grok 4.7 cost?
US$ 2 per million input tokens and US$ 6 per million output, for context up to 200,000 tokens. Above that the price doubles, to US$ 4 and US$ 12. For comparison: GPT-5.6 Sol charges US$ 4 and US$ 20, and Claude Fable 5.1 charges US$ 10 and US$ 50.
What does the asterisk on the 71% DeepSWE score mean?
It means that number was measured at a different effort setting from the rest of the column: "high" instead of "xhigh". Effort is how much the model thinks before answering, and it affects both result and cost. In practice, that row is not directly comparable with the others in the same column.
How do I know Grok 4.7 actually shipped?
By checking xAI documentation, not a screenshot. The identifier grok-4.7 is in the official model catalog with published pricing. That was the missing test during the weeks when the launch was announced and postponed several times.
Is it worth switching models because of these numbers?
Only after testing with your own task. A benchmark measures one specific set of problems, and the gap that shows up there may vanish in your case, or widen. The cheap path is running the same ten real tasks on both models and comparing result and cost before changing anything in production.
Can you trust a benchmark published by the vendor itself?
You can use it as a starting point, as long as you read the footnotes. In this table each model was measured at a different effort setting, one of the numbers carries a note saying it was measured at another setting, and whoever chose which seven tests to show is the party publishing. None of that makes the numbers false, but it makes the comparison less direct than it looks.
Rodrigo Fávaro
Founder of ROO3, a marketing and technology agency in São José do Rio Preto, Brazil. Builds AI products running in production (Tobia, gerar.app, Pense Mercado) and maintains the AI Benchmark, a public ranking of AI models. See ROO3 AI consulting.
X @rodmf LinkedIn rodrigofavaroKeep reading

Muse, Meta's AI agent: what changes for anyone who sells online
What Meta launched at Connect 2026 around its Muse agent, why Amazon blocked Muse from shopping, how the stock reacted...
14 min read
Opus 5.5: the expensive one got better and cheaper at the same time
The official Claude Opus 5.5 benchmarks against Fable 5.1, Opus 5 and GPT-6 Astra, the price per token, the 60% drop in...
11 min read
Amodei, Altman and Musk asked to slow AI down: what actually changes
Why Dario Amodei asked the industry to slow AI down in September 2026, why Altman and Musk agreed, the July incident...
12 min readWant to apply this in your company?
ROO3 diagnoses what can be automated first in your business. The first conversation is free.