Home › AI Benchmark

Which AI is best today? With a source for every number.

258 models from 32 labs, measured by the people who actually measure. No opinion here: only sources. How this is compiled

Latest release Sep 28, 2026

Claude Sonnet 5.5

Anthropic · API access

Still no ECI: the Epoch AI capability index is a composite and comes out in a later round after the release, so this model does not have a position in the intelligence ranking yet. What already exists is below: the individual Epoch AI tests, each next to the best result ever recorded on it, and, once published, the Arena and LiveBench scores. When the ECI comes out, the model moves into the table on its own.

Apex Agents 44.6% record 67.8% (Gemini 3.7 Flash)
CritPt 11.4% record 30.6% (GPT-5.5 Pro)
CursorBench 35.8% record 46.8% (Claude Fable 5.1)
FrontierCode 52.1% record 53.5% (Claude Fable 5)
FrontierSWE 61.9% record 65.5% (GPT-6 Astra)
ProofBench 100.0% record 100.0% (Claude Fable 5.1)
SciCode 49.1% record 61.0% (Claude Fable 5)
LiveBench, overall average 77.8 record 83.4 (Claude Fable 5.1)

Source of the scores: Epoch AI, under Creative Commons Attribution 4.0. ROO3 does not measure any model. The overall average is from LiveBench, under CC BY-SA 4.0. See every recent release without a score

Ranking by intelligence

ECI, from Epoch AI · 20 of 258 compiled on Sep 29, 2026 at 05:25
# Model Lab ECI ELO LiveBench Access Release
newGPT-6 Sol OpenAI no ECI yet no match 79.3 API access Sep 22, 2026
newGPT-6 Luna OpenAI no ECI yet no match 72.0 API access Sep 22, 2026
🥇 GPT-6 Astra OpenAI 166.60 no match 82.2 API access Sep 3, 2026
newClaude Opus 5.5 Anthropic no ECI yet 1,572 83.2 API access Sep 22, 2026
newClaude Sonnet 5.5 Anthropic no ECI yet no match 77.8 API access Sep 28, 2026
🥈 Claude Fable 5.1 Anthropic 165.00 no match 83.4 API access Sep 1, 2026
🥉 Claude Fable 5 Anthropic 163.60 1,546 83.0 API access Jun 9, 2026
4 Claude Opus 5 Anthropic 162.67 1,571 80.1 API access Jul 24, 2026
5 GPT-5.5 Pro OpenAI 162.45 no match no match API access Apr 23, 2026
6 GPT-5.6 Sol OpenAI 161.99 1,523 81.1 API access Jul 9, 2026
7 GPT-5.6 Terra OpenAI 159.31 1,514 77.9 API access Jul 9, 2026
8 GPT-5.5 OpenAI 159.25 1,535 80.2 API access Apr 23, 2026
9 GPT-5.4 Pro OpenAI 159.12 no match no match API access Mar 5, 2026
10 Claude Opus 4.8 Anthropic 158.30 1,509 76.2 API access May 28, 2026
11 Gemini 3.7 Flash Google DeepMind 157.72 1,553 78.8 API access Aug 13, 2026
12 Kimi K3 Moonshot 157.68 no match 79.2 Open weights (non-commercial) Jul 16, 2026
13 Gemini 3.8 Flash Google DeepMind 157.13 1,547 75.8 API access Sep 2, 2026
14 GPT-5.4 OpenAI 156.92 1,520 78.0 API access Mar 5, 2026
15 Muse Spark 1.3 Meta AI 156.89 no match 81.6 API access Sep 2, 2026
16 GPT-5.3 Codex OpenAI 156.84 no match no match API access Feb 5, 2026
17 Qwen 3.8 Max Alibaba 156.69 1,537 78.5 API access Aug 2, 2026
newGrok 4.7 xAI no ECI yet 1,476 77.4 API access Sep 21, 2026
18 Grok 4.6 xAI 156.48 1,479 78.0 API access Aug 12, 2026
19 Claude Opus 4.7 Anthropic 156.39 1,530 76.5 API access Apr 16, 2026
20 Claude Sonnet 5 Anthropic 156.34 1,491 76.0 API access Jun 30, 2026
Showing 20 of 258 models. The bar under the ECI starts at 155.11, not at zero: in this range a scale from zero would make them all look the same. A row marked new is a release Epoch AI has not scored yet. It appears next to its lab for comparison and does not take a position, because with no score there is no position. Compiled on Sep 29, 2026 at 05:25, Brasilia time.

ECI is capability measured on hard tests, compiled by Epoch AI. ELO is the preference of ordinary people in blind conversations, from LMArena. LiveBench, the average of 23 objective tasks with refreshed questions, is the third reading. The rankings do not agree, and the difference between them is the most useful information here. The default order is the ECI. None of these indexes is ours: we compile them, match the names and show where every number comes from.

258models in the ranking
32labs tracked
102with a human preference score
2sources, both CC BY 4.0
The frontier

Each dot is a model. The line is the record of its time.

Capability (ECI) against release date. The green line joins the models that, on the day they shipped, were the best that existed. It is the most direct read of how much the pace accelerated.

model record on that date Source: Epoch AI, CC BY 4.0.
Cost per task

How much each model delivers per dollar

Each line is a model, and each point is an effort level, from cheapest to most expensive. Higher means a better score; further right means a cheaper task. The curve shows what you gain by paying more for the same model, and where the difference stops being worth it.

CursorBench
Score on CursorBench by average cost per taskEach line is a model and each point is an effort level. The cost axis grows from right to left.20%25%30%35%40%45%50%55%60%$0$3$6$9$12$15$18Average cost per task · cheaper to the rightClaude Opus 5.5 · effort xhigh · 56.0% · $6.98 per task · 101,083 tokensClaude Opus 5.5 · effort max · 57.8% · $13.43 per task · 218,363 tokensClaude Sonnet 5.5 · effort low · 35.8% · $0.500 per task · 11,668 tokensClaude Sonnet 5.5 · effort medium · 39.2% · $0.700 per task · 16,036 tokensClaude Sonnet 5.5 · effort high · 47.8% · $1.67 per task · 37,391 tokensClaude Sonnet 5.5 · effort xhigh · 53.1% · $3.88 per task · 100,158 tokensClaude Sonnet 5.5 · effort max · 55.5% · $9.67 per task · 271,920 tokensClaude Fable 5.1 · effort low · 45.1% · $5.44 per task · 34,795 tokensClaude Fable 5.1 · effort medium · 46.8% · $7.05 per task · 45,411 tokensClaude Fable 5.1 · effort high · 49.2% · $9.08 per task · 58,438 tokensClaude Fable 5.1 · effort xhigh · 51.6% · $13.01 per task · 87,294 tokensClaude Fable 5.1 · effort max · 51.8% · $17.28 per task · 117,236 tokensClaude Opus 5 · effort low · 40.7% · $4.87 per task · 31,995 tokensClaude Opus 5 · effort medium · 43.3% · $6.94 per task · 45,272 tokensClaude Opus 5 · effort high · 44.7% · $9.00 per task · 61,405 tokensClaude Opus 5 · effort xhigh · 46.1% · $11.43 per task · 80,094 tokensClaude Opus 5 · effort max · 46.6% · $11.95 per task · 85,384 tokensGrok 4.7 · effort low · 33.1% · $1.58 per task · 15,677 tokensGrok 4.7 · effort medium · 41.6% · $3.49 per task · 36,683 tokensGrok 4.7 · effort high · 43.9% · $4.69 per task · 56,382 tokensGrok 4.7 · effort xhigh · 46.3% · $6.01 per task · 70,141 tokensGPT-5.6 Sol · effort low · 24.6% · $0.870 per task · 4,885 tokensGPT-5.6 Sol · effort medium · 31.1% · $1.77 per task · 10,111 tokensGPT-5.6 Sol · effort high · 35.7% · $2.85 per task · 16,174 tokensGPT-5.6 Sol · effort xhigh · 37.7% · $4.40 per task · 24,729 tokensGPT-5.6 Sol · effort max · 41.7% · $8.23 per task · 42,944 tokensMuse Spark 1.3 · effort minimal · 24.3% · $0.560 per task · 10,620 tokensMuse Spark 1.3 · effort low · 29.3% · $0.930 per task · 17,483 tokensMuse Spark 1.3 · effort medium · 32.6% · $1.49 per task · 27,255 tokensMuse Spark 1.3 · effort high · 33.4% · $1.66 per task · 30,654 tokensMuse Spark 1.3 · effort xhigh · 37.5% · $2.10 per task · 40,891 tokensMuse Spark 1.3 · effort max · 41.6% · $2.64 per task · 52,005 tokensGrok 4.6 · effort low · 33.4% · $2.25 per task · 16,307 tokensGrok 4.6 · effort medium · 36.1% · $3.48 per task · 24,893 tokensGrok 4.6 · effort high · 40.4% · $5.20 per task · 41,387 tokensGrok 4.6 · effort xhigh · 41.4% · $6.10 per task · 49,814 tokensGPT-5.6 Terra · effort low · 25.2% · $0.520 per task · 5,914 tokensGPT-5.6 Terra · effort medium · 27.6% · $0.640 per task · 7,307 tokensGPT-5.6 Terra · effort high · 30.7% · $1.11 per task · 13,162 tokensGPT-5.6 Terra · effort xhigh · 33.6% · $1.81 per task · 23,436 tokensGPT-5.6 Terra · effort max · 41.3% · $5.14 per task · 60,814 tokensGemini 3.8 Flash · effort medium · 37.3% · $4.06 per task · 128,364 tokensGemini 3.8 Flash · effort high · 39.6% · $4.70 per task · 162,565 tokensClaude Opus 5.5Claude Sonnet 5.5Claude Fable 5.1Claude Opus 5Grok 4.7GPT-5.6 SolMuse Spark 1.3Grok 4.6GPT-5.6 TerraGemini 3.8 Flash
Anthropic OpenAI Google DeepMind other labs

The 10 highest-scoring of the 12 models measured at more than one effort level (13 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0, from cursor.com/cursorbench.

See the numbers
ModelEffortScoreCost per taskTokens
Claude Opus 5.5max57.8%$13.43218,363
Claude Opus 5.5xhigh56.0%$6.98101,083
Claude Sonnet 5.5max55.5%$9.67271,920
Claude Sonnet 5.5xhigh53.1%$3.88100,158
Claude Sonnet 5.5high47.8%$1.6737,391
Claude Sonnet 5.5medium39.2%$0.70016,036
Claude Sonnet 5.5low35.8%$0.50011,668
Claude Fable 5.1max51.8%$17.28117,236
Claude Fable 5.1xhigh51.6%$13.0187,294
Claude Fable 5.1high49.2%$9.0858,438
Claude Fable 5.1medium46.8%$7.0545,411
Claude Fable 5.1low45.1%$5.4434,795
Claude Opus 5max46.6%$11.9585,384
Claude Opus 5xhigh46.1%$11.4380,094
Claude Opus 5high44.7%$9.0061,405
Claude Opus 5medium43.3%$6.9445,272
Claude Opus 5low40.7%$4.8731,995
Grok 4.7xhigh46.3%$6.0170,141
Grok 4.7high43.9%$4.6956,382
Grok 4.7medium41.6%$3.4936,683
Grok 4.7low33.1%$1.5815,677
GPT-5.6 Solmax41.7%$8.2342,944
GPT-5.6 Solxhigh37.7%$4.4024,729
GPT-5.6 Solhigh35.7%$2.8516,174
GPT-5.6 Solmedium31.1%$1.7710,111
GPT-5.6 Sollow24.6%$0.8704,885
Muse Spark 1.3max41.6%$2.6452,005
Muse Spark 1.3xhigh37.5%$2.1040,891
Muse Spark 1.3high33.4%$1.6630,654
Muse Spark 1.3medium32.6%$1.4927,255
Muse Spark 1.3low29.3%$0.93017,483
Muse Spark 1.3minimal24.3%$0.56010,620
Grok 4.6xhigh41.4%$6.1049,814
Grok 4.6high40.4%$5.2041,387
Grok 4.6medium36.1%$3.4824,893
Grok 4.6low33.4%$2.2516,307
GPT-5.6 Terramax41.3%$5.1460,814
GPT-5.6 Terraxhigh33.6%$1.8123,436
GPT-5.6 Terrahigh30.7%$1.1113,162
GPT-5.6 Terramedium27.6%$0.6407,307
GPT-5.6 Terralow25.2%$0.5205,914
Gemini 3.8 Flashhigh39.6%$4.70162,565
Gemini 3.8 Flashmedium37.3%$4.06128,364
DeepSWE
Score on DeepSWE by average cost per taskEach line is a model and each point is an effort level. The cost axis grows from right to left.0%10%20%30%40%50%60%70%80%$0$5$10$15$20$25Average cost per task · cheaper to the rightGPT-6 Astra · effort low · 67.0% · $2.19 per task · 10,580 tokensGPT-6 Astra · effort medium · 72.8% · $4.38 per task · 20,362 tokensGPT-6 Astra · effort high · 73.2% · $5.72 per task · 26,506 tokensGPT-6 Astra · effort xhigh · 74.1% · $6.52 per task · 29,557 tokensGPT-6 Astra · effort max · 73.2% · $12.37 per task · 61,149 tokensGemini 3.8 Flash · effort medium · 71.0% · $1.97 per task · 124,684 tokensGemini 3.8 Flash · effort high · 73.8% · $2.36 per task · 143,243 tokensClaude Opus 5 · effort low · 58.1% · $1.66 per task · 19,884 tokensClaude Opus 5 · effort medium · 68.9% · $3.29 per task · 36,982 tokensClaude Opus 5 · effort high · 72.8% · $6.08 per task · 64,207 tokensClaude Opus 5 · effort xhigh · 73.2% · $9.07 per task · 91,672 tokensClaude Opus 5 · effort max · 73.6% · $11.84 per task · 117,566 tokensGPT-5.6 Sol · effort low · 45.4% · $1.07 per task · 10,579 tokensGPT-5.6 Sol · effort medium · 61.1% · $1.86 per task · 18,425 tokensGPT-5.6 Sol · effort high · 69.4% · $3.47 per task · 28,450 tokensGPT-5.6 Sol · effort xhigh · 70.7% · $4.70 per task · 40,745 tokensGPT-5.6 Sol · effort max · 72.7% · $8.39 per task · 60,014 tokensClaude Fable 5 · effort low · 59.6% · $3.76 per task · 25,243 tokensClaude Fable 5 · effort medium · 65.4% · $6.09 per task · 40,201 tokensClaude Fable 5 · effort high · 68.6% · $9.18 per task · 57,287 tokensClaude Fable 5 · effort xhigh · 69.9% · $13.41 per task · 80,352 tokensClaude Fable 5 · effort max · 69.7% · $21.63 per task · 118,593 tokensGPT-5.6 Terra · effort low · 24.1% · $0.428 per task · 8,572 tokensGPT-5.6 Terra · effort medium · 35.1% · $0.583 per task · 11,747 tokensGPT-5.6 Terra · effort high · 53.8% · $1.13 per task · 21,517 tokensGPT-5.6 Terra · effort xhigh · 60.2% · $2.13 per task · 39,617 tokensGPT-5.6 Terra · effort max · 69.6% · $4.95 per task · 71,939 tokensGrok 4.6 · effort low · 41.6% · $1.04 per task · 16,458 tokensGrok 4.6 · effort medium · 67.5% · $3.45 per task · 49,764 tokensGrok 4.6 · effort high · 65.2% · $4.38 per task · 61,161 tokensGrok 4.6 · effort xhigh · 66.7% · $5.50 per task · 71,404 tokensGPT-5.6 Luna · effort low · 1.5% · $0.072 per task · 3,128 tokensGPT-5.6 Luna · effort medium · 11.3% · $0.216 per task · 8,180 tokensGPT-5.6 Luna · effort high · 44.2% · $0.778 per task · 25,778 tokensGPT-5.6 Luna · effort xhigh · 56.9% · $1.54 per task · 44,678 tokensGPT-5.6 Luna · effort max · 67.2% · $3.03 per task · 73,400 tokensGPT-5.5 · effort low · 27.0% · $1.20 per task · 9,443 tokensGPT-5.5 · effort medium · 54.0% · $2.75 per task · 19,625 tokensGPT-5.5 · effort high · 64.4% · $5.10 per task · 31,159 tokensGPT-5.5 · effort xhigh · 67.0% · $7.23 per task · 46,295 tokensGemini 3.7 Flash · effort low · 53.8% · $1.83 per task · 73,365 tokensGemini 3.7 Flash · effort medium · 65.5% · $2.03 per task · 93,991 tokensGemini 3.7 Flash · effort high · 65.3% · $2.18 per task · 107,248 tokensGPT-6 AstraGemini 3.8 FlashClaude Opus 5GPT-5.6 SolClaude Fable 5GPT-5.6 TerraGrok 4.6GPT-5.6 LunaGPT-5.5Gemini 3.7 Flash
Anthropic OpenAI Google DeepMind other labs

The 10 highest-scoring of the 14 models measured at more than one effort level (26 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0, from deepswe.datacurve.ai.

See the numbers
ModelEffortScoreCost per taskTokens
GPT-6 Astramax73.2%$12.3761,149
GPT-6 Astraxhigh74.1%$6.5229,557
GPT-6 Astrahigh73.2%$5.7226,506
GPT-6 Astramedium72.8%$4.3820,362
GPT-6 Astralow67.0%$2.1910,580
Gemini 3.8 Flashhigh73.8%$2.36143,243
Gemini 3.8 Flashmedium71.0%$1.97124,684
Claude Opus 5max73.6%$11.84117,566
Claude Opus 5xhigh73.2%$9.0791,672
Claude Opus 5high72.8%$6.0864,207
Claude Opus 5medium68.9%$3.2936,982
Claude Opus 5low58.1%$1.6619,884
GPT-5.6 Solmax72.7%$8.3960,014
GPT-5.6 Solxhigh70.7%$4.7040,745
GPT-5.6 Solhigh69.4%$3.4728,450
GPT-5.6 Solmedium61.1%$1.8618,425
GPT-5.6 Sollow45.4%$1.0710,579
Claude Fable 5max69.7%$21.63118,593
Claude Fable 5xhigh69.9%$13.4180,352
Claude Fable 5high68.6%$9.1857,287
Claude Fable 5medium65.4%$6.0940,201
Claude Fable 5low59.6%$3.7625,243
GPT-5.6 Terramax69.6%$4.9571,939
GPT-5.6 Terraxhigh60.2%$2.1339,617
GPT-5.6 Terrahigh53.8%$1.1321,517
GPT-5.6 Terramedium35.1%$0.58311,747
GPT-5.6 Terralow24.1%$0.4288,572
Grok 4.6xhigh66.7%$5.5071,404
Grok 4.6high65.2%$4.3861,161
Grok 4.6medium67.5%$3.4549,764
Grok 4.6low41.6%$1.0416,458
GPT-5.6 Lunamax67.2%$3.0373,400
GPT-5.6 Lunaxhigh56.9%$1.5444,678
GPT-5.6 Lunahigh44.2%$0.77825,778
GPT-5.6 Lunamedium11.3%$0.2168,180
GPT-5.6 Lunalow1.5%$0.0723,128
GPT-5.5xhigh67.0%$7.2346,295
GPT-5.5high64.4%$5.1031,159
GPT-5.5medium54.0%$2.7519,625
GPT-5.5low27.0%$1.209,443
Gemini 3.7 Flashhigh65.3%$2.18107,248
Gemini 3.7 Flashmedium65.5%$2.0393,991
Gemini 3.7 Flashlow53.8%$1.8373,365
ARC-AGI
Score on ARC-AGI by average cost per taskEach line is a model and each point is an effort level. The cost axis grows from right to left.45%50%55%60%65%70%75%80%85%90%95%100%$0.05$0.1$1$10$20Average cost per task (log scale) · cheaper to the rightGPT-5.6 Sol · effort low · 74.5% · $0.170 per taskGPT-5.6 Sol · effort medium · 92.5% · $0.220 per taskGPT-5.6 Sol · effort high · 97.0% · $0.300 per taskGPT-5.6 Sol · effort xhigh · 97.5% · $0.400 per taskGPT-5.6 Sol · effort max · 96.5% · $0.540 per taskGPT-5.5 Pro · effort xhigh · 95.0% · $4.52 per taskGPT-5.5 Pro · effort high · 96.5% · $4.53 per taskGPT-5.6 Terra · effort low · 60.2% · $0.090 per taskGPT-5.6 Terra · effort medium · 77.0% · $0.130 per taskGPT-5.6 Terra · effort high · 92.0% · $0.190 per taskGPT-5.6 Terra · effort xhigh · 94.0% · $0.260 per taskGPT-5.6 Terra · effort max · 96.5% · $0.550 per taskGPT-5.5 · effort low · 76.2% · $0.200 per taskGPT-5.5 · effort medium · 92.2% · $0.390 per taskGPT-5.5 · effort high · 94.5% · $0.560 per taskGPT-5.5 · effort xhigh · 95.0% · $0.730 per taskKimi K3 · effort low · 65.7% · $0.181 per taskKimi K3 · effort high · 86.7% · $0.480 per taskKimi K3 · effort max · 94.5% · $0.770 per taskGPT-5.4 · effort low · 68.2% · $0.150 per taskGPT-5.4 · effort medium · 86.2% · $0.250 per taskGPT-5.4 · effort high · 92.7% · $0.370 per taskGPT-5.4 · effort xhigh · 93.7% · $0.620 per taskClaude Opus 4.7 · effort low · 91.0% · $0.760 per taskClaude Opus 4.7 · effort medium · 91.0% · $1.04 per taskClaude Opus 4.7 · effort high · 93.5% · $1.41 per taskClaude Opus 4.7 · effort max · 92.0% · $2.58 per taskClaude Opus 4.8 · effort low · 88.0% · $0.671 per taskClaude Opus 4.8 · effort medium · 91.5% · $0.912 per taskClaude Opus 4.8 · effort high · 92.0% · $1.04 per taskClaude Opus 4.8 · effort max · 92.5% · $2.33 per taskGemini 3.5 Flash · effort minimal · 48.8% · $0.065 per taskGemini 3.5 Flash · effort high · 92.5% · $0.428 per taskGPT-5.2 Pro · effort medium · 81.2% · $3.98 per taskGPT-5.2 Pro · effort high · 85.7% · $5.87 per taskGPT-5.2 Pro · effort xhigh · 90.5% · $11.65 per taskGPT-5.6 SolGPT-5.5 ProGPT-5.6 TerraGPT-5.5Kimi K3GPT-5.4Claude Opus 4.7Claude Opus 4.8Gemini 3.5 FlashGPT-5.2 Pro
Anthropic OpenAI Google DeepMind other labs

The 10 highest-scoring of the 24 models measured at more than one effort level (92 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0, from arcprize.org/leaderboard.

See the numbers
ModelEffortScoreCost per taskTokens
GPT-5.6 Solmax96.5%$0.540-
GPT-5.6 Solxhigh97.5%$0.400-
GPT-5.6 Solhigh97.0%$0.300-
GPT-5.6 Solmedium92.5%$0.220-
GPT-5.6 Sollow74.5%$0.170-
GPT-5.5 Prohigh96.5%$4.53-
GPT-5.5 Proxhigh95.0%$4.52-
GPT-5.6 Terramax96.5%$0.550-
GPT-5.6 Terraxhigh94.0%$0.260-
GPT-5.6 Terrahigh92.0%$0.190-
GPT-5.6 Terramedium77.0%$0.130-
GPT-5.6 Terralow60.2%$0.090-
GPT-5.5xhigh95.0%$0.730-
GPT-5.5high94.5%$0.560-
GPT-5.5medium92.2%$0.390-
GPT-5.5low76.2%$0.200-
Kimi K3max94.5%$0.770-
Kimi K3high86.7%$0.480-
Kimi K3low65.7%$0.181-
GPT-5.4xhigh93.7%$0.620-
GPT-5.4high92.7%$0.370-
GPT-5.4medium86.2%$0.250-
GPT-5.4low68.2%$0.150-
Claude Opus 4.7max92.0%$2.58-
Claude Opus 4.7high93.5%$1.41-
Claude Opus 4.7medium91.0%$1.04-
Claude Opus 4.7low91.0%$0.760-
Claude Opus 4.8max92.5%$2.33-
Claude Opus 4.8high92.0%$1.04-
Claude Opus 4.8medium91.5%$0.912-
Claude Opus 4.8low88.0%$0.671-
Gemini 3.5 Flashhigh92.5%$0.428-
Gemini 3.5 Flashminimal48.8%$0.065-
GPT-5.2 Proxhigh90.5%$11.65-
GPT-5.2 Prohigh85.7%$5.87-
GPT-5.2 Promedium81.2%$3.98-
ARC-AGI 2
Score on ARC-AGI 2 by average cost per taskEach line is a model and each point is an effort level. The cost axis grows from right to left.0%10%20%30%40%50%60%70%80%90%100%$0.1$1$10$20Average cost per task (log scale) · cheaper to the rightGPT-5.6 Sol · effort low · 42.5% · $0.320 per taskGPT-5.6 Sol · effort medium · 67.1% · $0.470 per taskGPT-5.6 Sol · effort high · 85.4% · $0.740 per taskGPT-5.6 Sol · effort xhigh · 90.0% · $1.04 per taskGPT-5.6 Sol · effort max · 92.5% · $1.44 per taskGPT-5.5 · effort low · 33.3% · $0.350 per taskGPT-5.5 · effort medium · 70.4% · $0.860 per taskGPT-5.5 · effort high · 83.3% · $1.45 per taskGPT-5.5 · effort xhigh · 85.0% · $1.87 per taskGPT-5.5 Pro · effort high · 84.6% · $10.51 per taskGPT-5.5 Pro · effort xhigh · 84.2% · $10.76 per taskGPT-5.6 Terra · effort low · 18.8% · $0.170 per taskGPT-5.6 Terra · effort medium · 37.5% · $0.270 per taskGPT-5.6 Terra · effort high · 67.1% · $0.550 per taskGPT-5.6 Terra · effort xhigh · 74.2% · $0.690 per taskGPT-5.6 Terra · effort max · 83.9% · $1.09 per taskClaude Opus 4.7 · effort low · 62.1% · $2.38 per taskClaude Opus 4.7 · effort medium · 67.5% · $2.96 per taskClaude Opus 4.7 · effort high · 68.3% · $3.17 per taskClaude Opus 4.7 · effort max · 75.8% · $7.43 per taskGPT-5.4 · effort low · 29.2% · $0.270 per taskGPT-5.4 · effort medium · 55.4% · $0.680 per taskGPT-5.4 · effort high · 67.5% · $1.02 per taskGPT-5.4 · effort xhigh · 74.0% · $1.52 per taskClaude Opus 4.8 · effort low · 62.2% · $1.68 per taskClaude Opus 4.8 · effort medium · 71.7% · $2.39 per taskClaude Opus 4.8 · effort high · 72.1% · $2.74 per taskGemini 3.5 Flash · effort minimal · 8.9% · $0.107 per taskGemini 3.5 Flash · effort high · 72.1% · $0.850 per taskClaude Sonnet 4.6 · effort high · 60.4% · $2.70 per taskClaude Sonnet 4.6 · effort max · 58.3% · $2.72 per taskKimi K3 · effort low · 12.4% · $0.254 per taskKimi K3 · effort high · 55.0% · $0.947 per taskKimi K3 · effort max · 60.4% · $1.59 per taskGPT-5.6 SolGPT-5.5GPT-5.5 ProGPT-5.6 TerraClaude Opus 4.7GPT-5.4Claude Opus 4.8Gemini 3.5 FlashClaude Sonnet 4.6Kimi K3
Anthropic OpenAI Google DeepMind other labs

The 10 highest-scoring of the 24 models measured at more than one effort level (89 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0.

See the numbers
ModelEffortScoreCost per taskTokens
GPT-5.6 Solmax92.5%$1.44-
GPT-5.6 Solxhigh90.0%$1.04-
GPT-5.6 Solhigh85.4%$0.740-
GPT-5.6 Solmedium67.1%$0.470-
GPT-5.6 Sollow42.5%$0.320-
GPT-5.5xhigh85.0%$1.87-
GPT-5.5high83.3%$1.45-
GPT-5.5medium70.4%$0.860-
GPT-5.5low33.3%$0.350-
GPT-5.5 Proxhigh84.2%$10.76-
GPT-5.5 Prohigh84.6%$10.51-
GPT-5.6 Terramax83.9%$1.09-
GPT-5.6 Terraxhigh74.2%$0.690-
GPT-5.6 Terrahigh67.1%$0.550-
GPT-5.6 Terramedium37.5%$0.270-
GPT-5.6 Terralow18.8%$0.170-
Claude Opus 4.7max75.8%$7.43-
Claude Opus 4.7high68.3%$3.17-
Claude Opus 4.7medium67.5%$2.96-
Claude Opus 4.7low62.1%$2.38-
GPT-5.4xhigh74.0%$1.52-
GPT-5.4high67.5%$1.02-
GPT-5.4medium55.4%$0.680-
GPT-5.4low29.2%$0.270-
Claude Opus 4.8high72.1%$2.74-
Claude Opus 4.8medium71.7%$2.39-
Claude Opus 4.8low62.2%$1.68-
Gemini 3.5 Flashhigh72.1%$0.850-
Gemini 3.5 Flashminimal8.9%$0.107-
Claude Sonnet 4.6max58.3%$2.72-
Claude Sonnet 4.6high60.4%$2.70-
Kimi K3max60.4%$1.59-
Kimi K3high55.0%$0.947-
Kimi K3low12.4%$0.254-
Who is ahead

The best model from each lab

Ordered by the ECI of the most capable model each organization has in the ranking.

OpenAI

GPT-6 Astra
166.60
36 models in the ranking · United States of America

Anthropic

Claude Fable 5.1
165.00
23 models in the ranking · United States of America

Google DeepMind

Gemini 3.7 Flash
157.72
27 models in the ranking · United States of America

Moonshot

Kimi K3
157.68
7 models in the ranking · China

Meta AI

Muse Spark 1.3
156.89
22 models in the ranking · United States of America

Alibaba

Qwen 3.8 Max
156.69
36 models in the ranking · China

xAI

Grok 4.6
156.48
10 models in the ranking · United States of America

Z.ai (Zhipu AI)

GLM-5.3
155.56
6 models in the ranking · China

DeepSeek

DeepSeek V4 Pro 0813
155.39
15 models in the ranking · China

Thinking Machines

Inkling-Small
150.13
2 models in the ranking · United States of America

Xiaomi Corp

MiMo-V2.5-Pro
149.27
1 model in the ranking · China

MiniMax

MiniMax-M3
147.00
3 models in the ranking · China
Measurement queue

Recent releases, still without a score

Epoch AI registers the model on the day it comes out and computes the capability score later. In that window the model exists, may already be in everyone's hands, and has no ECI yet. It does not take a position in the ranking above, which is ordered by capability, and it is listed here. When the Arena and LiveBench have already measured it, their scores appear; when not, the column stays empty, as in the rest of the page. Today there are 12 models in that situation. The capability column says how many individual Epoch AI tests already have a published result for that model: it is genuine measurement, just not yet consolidated into the index.

Model Lab Release ELO LiveBench Capability Access
Claude Sonnet 5.5 Anthropic Sep 28, 2026 no match 77.8 7 tests already measured API access
Claude Opus 5.5 Anthropic Sep 22, 2026 1,572 83.2 14 tests already measured API access
GPT-6 Luna OpenAI Sep 22, 2026 no match 72.0 12 tests already measured API access
GPT-6 Sol OpenAI Sep 22, 2026 no match 79.3 13 tests already measured API access
Grok 4.7 xAI Sep 21, 2026 1,476 77.4 6 tests already measured API access
Step 5 Preview StepFun Sep 18, 2026 no match no match 2 tests already measured API access
SWE-2 Cognition Sep 10, 2026 no match no match 1 test already measured Hosted access (no API)
Mercury 2.5 Inception Labs Sep 8, 2026 no match no match 2 tests already measured API access
Tencent Hy4 preview Tencent Aug 28, 2026 no match no match 1 test already measured Open weights (unrestricted)
Qwen3.8 Flash Alibaba Aug 26, 2026 no match no match 1 test already measured API access
Grok Build 0.1 xAI May 29, 2026 no match 67.8 awaiting measurement API access
GPT-5.2 Codex OpenAI Dec 18, 2025 no match 74.0 awaiting measurement API access
The 12 most recent releases without a capability score. As soon as Epoch AI publishes the ECI, the model enters the ranking by itself.
How this is compiled

ROO3 does not measure any model

This page has no lab and runs no benchmark. It gathers, translates and credits two public measurements made by people who have the infrastructure for it. Saying otherwise would be the first lie, and after it no number here would be worth anything.

We order by ECI, not by an average

ECI, ELO and LiveBench are on different scales: the ECI sits around 150, the ELO around 1400 and LiveBench goes from 0 to 100. Adding them into a "score of our own" would produce a nice-looking number that means nothing. The ranking follows the ECI, and the other scores sit alongside for you to compare with your own eyes.

An empty column is an answer, not a failure

The sources write the same model in different ways. Matching by text similarity is exactly how an aggregator starts lying without throwing an error: all it takes is pasting a "Flash" ELO onto the "Pro" with the same name. Here the match is exact or it does not exist. Today, 102 of the 258 models have a confirmed match on the Arena and 48 on LiveBench.

One model, its ceiling

The same model shows up in the sources several times, once per reasoning-effort setting. Instead of repeating the top of the ranking with the same name five times, we keep the best result of each family in each column. It is that model's real ceiling.

A source being down does not empty a column

When one of the sources does not answer, the sync keeps the previous day number and logs the failure. The opposite, silently clearing the column, is the kind of defect nobody notices: the page stays up, looking fine, with one piece of data missing.

Only sources that allow redistribution

There are more famous rankings than these. They were left out because their terms of use forbid republishing the data on a third-party commercial site. We prefer sources that allow it in writing to another one that would bring more traffic and a legal problem.

Six times a day, automatic

A scheduled job downloads the three sources, checks that each license is still the same (CC BY for Epoch AI and the Arena, CC BY-SA for LiveBench), keeps the fingerprint of the day's file and saves. If a license changes, syncing that source stops on its own and nothing new from it is saved.

Sources and credits

Whose data this is

Epoch AI and the Arena are under Creative Commons Attribution 4.0, and LiveBench under Attribution-ShareAlike 4.0. All three allow commercial use and redistribution as long as they are credited. The LiveBench numbers stay under the same CC BY-SA. The credit below is generated from the update's own record, not written by hand, so it does not get lost in a future edit.

Epoch AI, "AI Benchmarking Hub". Published at epoch.ai. Data under the CC BY 4.0 licence, reordered by roo3.co.

Epoch AI, AI Benchmarking Hub · CC BY 4.0

253 records read on Sep 29, 2026 at 05:25.

The Arena (LMArena) leaderboard, from the public lmarena-ai/leaderboard-dataset on Hugging Face, under the CC BY 4.0 licence. We are not affiliated with the Arena.

Arena (LMArena), leaderboard-dataset · CC BY 4.0

126 records read on Sep 29, 2026 at 05:25.

LiveBench, by the LiveBench team, published at livebench.ai under the CC BY-SA 4.0 license. The overall average is recomputed by roo3.co with the same formula as the website (average of the categories), and those numbers stay under CC BY-SA 4.0. We are not affiliated with LiveBench.

LiveBench (livebench.ai) · CC BY-SA 4.0

64 records read on Sep 29, 2026 at 05:25.

ROO3 is not affiliated with Epoch AI, the Arena (LMArena) or LiveBench. The data was compiled, ordered and translated into Portuguese; the values were not changed, and the LiveBench overall average is recomputed with the same formula as its website. Trademarks belong to their respective owners.

Frequently asked questions

What these numbers mean

Which AI is best today?

It depends on what you mean by best. For capability measured on hard tests, the top of the ECI. For the preference of ordinary people in everyday conversations, the top of the ELO. LiveBench is a third reading, with refreshed questions to reduce the chance that the model has already seen the test. The rankings do not agree, which is why this page shows the columns side by side instead of inventing a single score.

What is the ECI?

Epoch Capabilities Index, from Epoch AI. A composite index that combines a model performance across several hard benchmarks into a single scale, which allows comparing models that never took exactly the same tests. The higher, the better.

What is the Arena ELO?

It is the Arena human preference score. People send the same question to two models without knowing which they are, pick the better answer, and the system calculates a score on the same rating scheme as chess. It measures what people like to receive, which is not the same as raw capability.

What is LiveBench?

It is a benchmark with 23 objective tasks in 7 categories (reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following), with questions refreshed every six months to reduce the chance that the model has already seen the test. The score goes from 0 to 100 and is the average of the categories, the same formula the LiveBench website uses. It usually publishes a release's score before the Epoch AI index.

Why do some models have an empty ELO column?

Because that model was not found in the Arena with certainty. Matching names by similarity is how a ranking starts lying without throwing any error. When there is no exact match, the column stays empty. We prefer the gap to the wrong number.

I cannot find a model that just shipped. Where is it?

If it came out recently, it shows up in the release card at the top of the page and in the "Recent releases, not scored yet" list, right below the table. Epoch AI registers the model on release day and computes the capability score later, so there is a window when the model is public but has no ECI yet. In that window it does not take a position in the ranking, which is ordered by capability, but it appears with its LiveBench and Arena scores when they already exist, which usually happens before the ECI. We do not estimate the missing score.

Do open weight models appear here?

They do, and you can filter for them alone in the bar above the table. The Access column shows whether the model is served through an API, whether the weights are open or whether access is restricted, following Epoch AI own classification.

Can I use this data in my own work?

The original data comes from the three sources credited below: Epoch AI and Arena under Creative Commons Attribution 4.0, and LiveBench under Attribution-ShareAlike 4.0. You can use it, commercially too, as long as you credit the original sources the same way we do here; anything taken from LiveBench stays under the same CC BY-SA. The credit is not a courtesy: it is the condition of the license.

How often does it update?

Six times a day, automatically. The date and time of the last update are at the top of the page and in the table footer. If one of the sources is down, the page keeps the numbers from the previous update instead of blanking the column.

Knowing which AI is best is easy. Using it in your business is the hard part.

We roll AI out in real companies: customer service, content, process. With numbers at the end of the month, not with promises.

Talk about AI in my company
Quick reply on WhatsApp. São José do Rio Preto, Brazil, and remote everywhere.