Which AI is best today? With a source for every number.
258 models from 32 labs, measured by the people who actually measure. No opinion here: only sources. How this is compiled
Claude Sonnet 5.5
Anthropic · API access
Still no ECI: the Epoch AI capability index is a composite and comes out in a later round after the release, so this model does not have a position in the intelligence ranking yet. What already exists is below: the individual Epoch AI tests, each next to the best result ever recorded on it, and, once published, the Arena and LiveBench scores. When the ECI comes out, the model moves into the table on its own.
Source of the scores: Epoch AI, under Creative Commons Attribution 4.0. ROO3 does not measure any model. The overall average is from LiveBench, under CC BY-SA 4.0. See every recent release without a score
🏆 Ranking by intelligence
ECI, from Epoch AI · 20 of 258 compiled on Sep 29, 2026 at 05:25| # | Model | Lab | ECI | ELO | LiveBench | Access | Release |
|---|---|---|---|---|---|---|---|
| newGPT-6 Sol | OpenAI | no ECI yet | no match | 79.3 | API access | Sep 22, 2026 | |
| newGPT-6 Luna | OpenAI | no ECI yet | no match | 72.0 | API access | Sep 22, 2026 | |
| 🥇 | GPT-6 Astra | OpenAI | 166.60 | no match | 82.2 | API access | Sep 3, 2026 |
| newClaude Opus 5.5 | Anthropic | no ECI yet | 1,572 | 83.2 | API access | Sep 22, 2026 | |
| newClaude Sonnet 5.5 | Anthropic | no ECI yet | no match | 77.8 | API access | Sep 28, 2026 | |
| 🥈 | Claude Fable 5.1 | Anthropic | 165.00 | no match | 83.4 | API access | Sep 1, 2026 |
| 🥉 | Claude Fable 5 | Anthropic | 163.60 | 1,546 | 83.0 | API access | Jun 9, 2026 |
| 4 | Claude Opus 5 | Anthropic | 162.67 | 1,571 | 80.1 | API access | Jul 24, 2026 |
| 5 | GPT-5.5 Pro | OpenAI | 162.45 | no match | no match | API access | Apr 23, 2026 |
| 6 | GPT-5.6 Sol | OpenAI | 161.99 | 1,523 | 81.1 | API access | Jul 9, 2026 |
| 7 | GPT-5.6 Terra | OpenAI | 159.31 | 1,514 | 77.9 | API access | Jul 9, 2026 |
| 8 | GPT-5.5 | OpenAI | 159.25 | 1,535 | 80.2 | API access | Apr 23, 2026 |
| 9 | GPT-5.4 Pro | OpenAI | 159.12 | no match | no match | API access | Mar 5, 2026 |
| 10 | Claude Opus 4.8 | Anthropic | 158.30 | 1,509 | 76.2 | API access | May 28, 2026 |
| 11 | Gemini 3.7 Flash | Google DeepMind | 157.72 | 1,553 | 78.8 | API access | Aug 13, 2026 |
| 12 | Kimi K3 | Moonshot | 157.68 | no match | 79.2 | Open weights (non-commercial) | Jul 16, 2026 |
| 13 | Gemini 3.8 Flash | Google DeepMind | 157.13 | 1,547 | 75.8 | API access | Sep 2, 2026 |
| 14 | GPT-5.4 | OpenAI | 156.92 | 1,520 | 78.0 | API access | Mar 5, 2026 |
| 15 | Muse Spark 1.3 | Meta AI | 156.89 | no match | 81.6 | API access | Sep 2, 2026 |
| 16 | GPT-5.3 Codex | OpenAI | 156.84 | no match | no match | API access | Feb 5, 2026 |
| 17 | Qwen 3.8 Max | Alibaba | 156.69 | 1,537 | 78.5 | API access | Aug 2, 2026 |
| newGrok 4.7 | xAI | no ECI yet | 1,476 | 77.4 | API access | Sep 21, 2026 | |
| 18 | Grok 4.6 | xAI | 156.48 | 1,479 | 78.0 | API access | Aug 12, 2026 |
| 19 | Claude Opus 4.7 | Anthropic | 156.39 | 1,530 | 76.5 | API access | Apr 16, 2026 |
| 20 | Claude Sonnet 5 | Anthropic | 156.34 | 1,491 | 76.0 | API access | Jun 30, 2026 |
ECI is capability measured on hard tests, compiled by Epoch AI. ELO is the preference of ordinary people in blind conversations, from LMArena. LiveBench, the average of 23 objective tasks with refreshed questions, is the third reading. The rankings do not agree, and the difference between them is the most useful information here. The default order is the ECI. None of these indexes is ours: we compile them, match the names and show where every number comes from.
Each dot is a model. The line is the record of its time.
Capability (ECI) against release date. The green line joins the models that, on the day they shipped, were the best that existed. It is the most direct read of how much the pace accelerated.
How much each model delivers per dollar
Each line is a model, and each point is an effort level, from cheapest to most expensive. Higher means a better score; further right means a cheaper task. The curve shows what you gain by paying more for the same model, and where the difference stops being worth it.
The 10 highest-scoring of the 12 models measured at more than one effort level (13 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0, from cursor.com/cursorbench.
See the numbers
| Model | Effort | Score | Cost per task | Tokens |
|---|---|---|---|---|
| Claude Opus 5.5 | max | 57.8% | $13.43 | 218,363 |
| Claude Opus 5.5 | xhigh | 56.0% | $6.98 | 101,083 |
| Claude Sonnet 5.5 | max | 55.5% | $9.67 | 271,920 |
| Claude Sonnet 5.5 | xhigh | 53.1% | $3.88 | 100,158 |
| Claude Sonnet 5.5 | high | 47.8% | $1.67 | 37,391 |
| Claude Sonnet 5.5 | medium | 39.2% | $0.700 | 16,036 |
| Claude Sonnet 5.5 | low | 35.8% | $0.500 | 11,668 |
| Claude Fable 5.1 | max | 51.8% | $17.28 | 117,236 |
| Claude Fable 5.1 | xhigh | 51.6% | $13.01 | 87,294 |
| Claude Fable 5.1 | high | 49.2% | $9.08 | 58,438 |
| Claude Fable 5.1 | medium | 46.8% | $7.05 | 45,411 |
| Claude Fable 5.1 | low | 45.1% | $5.44 | 34,795 |
| Claude Opus 5 | max | 46.6% | $11.95 | 85,384 |
| Claude Opus 5 | xhigh | 46.1% | $11.43 | 80,094 |
| Claude Opus 5 | high | 44.7% | $9.00 | 61,405 |
| Claude Opus 5 | medium | 43.3% | $6.94 | 45,272 |
| Claude Opus 5 | low | 40.7% | $4.87 | 31,995 |
| Grok 4.7 | xhigh | 46.3% | $6.01 | 70,141 |
| Grok 4.7 | high | 43.9% | $4.69 | 56,382 |
| Grok 4.7 | medium | 41.6% | $3.49 | 36,683 |
| Grok 4.7 | low | 33.1% | $1.58 | 15,677 |
| GPT-5.6 Sol | max | 41.7% | $8.23 | 42,944 |
| GPT-5.6 Sol | xhigh | 37.7% | $4.40 | 24,729 |
| GPT-5.6 Sol | high | 35.7% | $2.85 | 16,174 |
| GPT-5.6 Sol | medium | 31.1% | $1.77 | 10,111 |
| GPT-5.6 Sol | low | 24.6% | $0.870 | 4,885 |
| Muse Spark 1.3 | max | 41.6% | $2.64 | 52,005 |
| Muse Spark 1.3 | xhigh | 37.5% | $2.10 | 40,891 |
| Muse Spark 1.3 | high | 33.4% | $1.66 | 30,654 |
| Muse Spark 1.3 | medium | 32.6% | $1.49 | 27,255 |
| Muse Spark 1.3 | low | 29.3% | $0.930 | 17,483 |
| Muse Spark 1.3 | minimal | 24.3% | $0.560 | 10,620 |
| Grok 4.6 | xhigh | 41.4% | $6.10 | 49,814 |
| Grok 4.6 | high | 40.4% | $5.20 | 41,387 |
| Grok 4.6 | medium | 36.1% | $3.48 | 24,893 |
| Grok 4.6 | low | 33.4% | $2.25 | 16,307 |
| GPT-5.6 Terra | max | 41.3% | $5.14 | 60,814 |
| GPT-5.6 Terra | xhigh | 33.6% | $1.81 | 23,436 |
| GPT-5.6 Terra | high | 30.7% | $1.11 | 13,162 |
| GPT-5.6 Terra | medium | 27.6% | $0.640 | 7,307 |
| GPT-5.6 Terra | low | 25.2% | $0.520 | 5,914 |
| Gemini 3.8 Flash | high | 39.6% | $4.70 | 162,565 |
| Gemini 3.8 Flash | medium | 37.3% | $4.06 | 128,364 |
The 10 highest-scoring of the 14 models measured at more than one effort level (26 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0, from deepswe.datacurve.ai.
See the numbers
| Model | Effort | Score | Cost per task | Tokens |
|---|---|---|---|---|
| GPT-6 Astra | max | 73.2% | $12.37 | 61,149 |
| GPT-6 Astra | xhigh | 74.1% | $6.52 | 29,557 |
| GPT-6 Astra | high | 73.2% | $5.72 | 26,506 |
| GPT-6 Astra | medium | 72.8% | $4.38 | 20,362 |
| GPT-6 Astra | low | 67.0% | $2.19 | 10,580 |
| Gemini 3.8 Flash | high | 73.8% | $2.36 | 143,243 |
| Gemini 3.8 Flash | medium | 71.0% | $1.97 | 124,684 |
| Claude Opus 5 | max | 73.6% | $11.84 | 117,566 |
| Claude Opus 5 | xhigh | 73.2% | $9.07 | 91,672 |
| Claude Opus 5 | high | 72.8% | $6.08 | 64,207 |
| Claude Opus 5 | medium | 68.9% | $3.29 | 36,982 |
| Claude Opus 5 | low | 58.1% | $1.66 | 19,884 |
| GPT-5.6 Sol | max | 72.7% | $8.39 | 60,014 |
| GPT-5.6 Sol | xhigh | 70.7% | $4.70 | 40,745 |
| GPT-5.6 Sol | high | 69.4% | $3.47 | 28,450 |
| GPT-5.6 Sol | medium | 61.1% | $1.86 | 18,425 |
| GPT-5.6 Sol | low | 45.4% | $1.07 | 10,579 |
| Claude Fable 5 | max | 69.7% | $21.63 | 118,593 |
| Claude Fable 5 | xhigh | 69.9% | $13.41 | 80,352 |
| Claude Fable 5 | high | 68.6% | $9.18 | 57,287 |
| Claude Fable 5 | medium | 65.4% | $6.09 | 40,201 |
| Claude Fable 5 | low | 59.6% | $3.76 | 25,243 |
| GPT-5.6 Terra | max | 69.6% | $4.95 | 71,939 |
| GPT-5.6 Terra | xhigh | 60.2% | $2.13 | 39,617 |
| GPT-5.6 Terra | high | 53.8% | $1.13 | 21,517 |
| GPT-5.6 Terra | medium | 35.1% | $0.583 | 11,747 |
| GPT-5.6 Terra | low | 24.1% | $0.428 | 8,572 |
| Grok 4.6 | xhigh | 66.7% | $5.50 | 71,404 |
| Grok 4.6 | high | 65.2% | $4.38 | 61,161 |
| Grok 4.6 | medium | 67.5% | $3.45 | 49,764 |
| Grok 4.6 | low | 41.6% | $1.04 | 16,458 |
| GPT-5.6 Luna | max | 67.2% | $3.03 | 73,400 |
| GPT-5.6 Luna | xhigh | 56.9% | $1.54 | 44,678 |
| GPT-5.6 Luna | high | 44.2% | $0.778 | 25,778 |
| GPT-5.6 Luna | medium | 11.3% | $0.216 | 8,180 |
| GPT-5.6 Luna | low | 1.5% | $0.072 | 3,128 |
| GPT-5.5 | xhigh | 67.0% | $7.23 | 46,295 |
| GPT-5.5 | high | 64.4% | $5.10 | 31,159 |
| GPT-5.5 | medium | 54.0% | $2.75 | 19,625 |
| GPT-5.5 | low | 27.0% | $1.20 | 9,443 |
| Gemini 3.7 Flash | high | 65.3% | $2.18 | 107,248 |
| Gemini 3.7 Flash | medium | 65.5% | $2.03 | 93,991 |
| Gemini 3.7 Flash | low | 53.8% | $1.83 | 73,365 |
The 10 highest-scoring of the 24 models measured at more than one effort level (92 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0, from arcprize.org/leaderboard.
See the numbers
| Model | Effort | Score | Cost per task | Tokens |
|---|---|---|---|---|
| GPT-5.6 Sol | max | 96.5% | $0.540 | - |
| GPT-5.6 Sol | xhigh | 97.5% | $0.400 | - |
| GPT-5.6 Sol | high | 97.0% | $0.300 | - |
| GPT-5.6 Sol | medium | 92.5% | $0.220 | - |
| GPT-5.6 Sol | low | 74.5% | $0.170 | - |
| GPT-5.5 Pro | high | 96.5% | $4.53 | - |
| GPT-5.5 Pro | xhigh | 95.0% | $4.52 | - |
| GPT-5.6 Terra | max | 96.5% | $0.550 | - |
| GPT-5.6 Terra | xhigh | 94.0% | $0.260 | - |
| GPT-5.6 Terra | high | 92.0% | $0.190 | - |
| GPT-5.6 Terra | medium | 77.0% | $0.130 | - |
| GPT-5.6 Terra | low | 60.2% | $0.090 | - |
| GPT-5.5 | xhigh | 95.0% | $0.730 | - |
| GPT-5.5 | high | 94.5% | $0.560 | - |
| GPT-5.5 | medium | 92.2% | $0.390 | - |
| GPT-5.5 | low | 76.2% | $0.200 | - |
| Kimi K3 | max | 94.5% | $0.770 | - |
| Kimi K3 | high | 86.7% | $0.480 | - |
| Kimi K3 | low | 65.7% | $0.181 | - |
| GPT-5.4 | xhigh | 93.7% | $0.620 | - |
| GPT-5.4 | high | 92.7% | $0.370 | - |
| GPT-5.4 | medium | 86.2% | $0.250 | - |
| GPT-5.4 | low | 68.2% | $0.150 | - |
| Claude Opus 4.7 | max | 92.0% | $2.58 | - |
| Claude Opus 4.7 | high | 93.5% | $1.41 | - |
| Claude Opus 4.7 | medium | 91.0% | $1.04 | - |
| Claude Opus 4.7 | low | 91.0% | $0.760 | - |
| Claude Opus 4.8 | max | 92.5% | $2.33 | - |
| Claude Opus 4.8 | high | 92.0% | $1.04 | - |
| Claude Opus 4.8 | medium | 91.5% | $0.912 | - |
| Claude Opus 4.8 | low | 88.0% | $0.671 | - |
| Gemini 3.5 Flash | high | 92.5% | $0.428 | - |
| Gemini 3.5 Flash | minimal | 48.8% | $0.065 | - |
| GPT-5.2 Pro | xhigh | 90.5% | $11.65 | - |
| GPT-5.2 Pro | high | 85.7% | $5.87 | - |
| GPT-5.2 Pro | medium | 81.2% | $3.98 | - |
The 10 highest-scoring of the 24 models measured at more than one effort level (89 measured in total on this benchmark). A model released in the last few days is added once the source publishes its measurement. Source: Epoch AI, CC BY 4.0.
See the numbers
| Model | Effort | Score | Cost per task | Tokens |
|---|---|---|---|---|
| GPT-5.6 Sol | max | 92.5% | $1.44 | - |
| GPT-5.6 Sol | xhigh | 90.0% | $1.04 | - |
| GPT-5.6 Sol | high | 85.4% | $0.740 | - |
| GPT-5.6 Sol | medium | 67.1% | $0.470 | - |
| GPT-5.6 Sol | low | 42.5% | $0.320 | - |
| GPT-5.5 | xhigh | 85.0% | $1.87 | - |
| GPT-5.5 | high | 83.3% | $1.45 | - |
| GPT-5.5 | medium | 70.4% | $0.860 | - |
| GPT-5.5 | low | 33.3% | $0.350 | - |
| GPT-5.5 Pro | xhigh | 84.2% | $10.76 | - |
| GPT-5.5 Pro | high | 84.6% | $10.51 | - |
| GPT-5.6 Terra | max | 83.9% | $1.09 | - |
| GPT-5.6 Terra | xhigh | 74.2% | $0.690 | - |
| GPT-5.6 Terra | high | 67.1% | $0.550 | - |
| GPT-5.6 Terra | medium | 37.5% | $0.270 | - |
| GPT-5.6 Terra | low | 18.8% | $0.170 | - |
| Claude Opus 4.7 | max | 75.8% | $7.43 | - |
| Claude Opus 4.7 | high | 68.3% | $3.17 | - |
| Claude Opus 4.7 | medium | 67.5% | $2.96 | - |
| Claude Opus 4.7 | low | 62.1% | $2.38 | - |
| GPT-5.4 | xhigh | 74.0% | $1.52 | - |
| GPT-5.4 | high | 67.5% | $1.02 | - |
| GPT-5.4 | medium | 55.4% | $0.680 | - |
| GPT-5.4 | low | 29.2% | $0.270 | - |
| Claude Opus 4.8 | high | 72.1% | $2.74 | - |
| Claude Opus 4.8 | medium | 71.7% | $2.39 | - |
| Claude Opus 4.8 | low | 62.2% | $1.68 | - |
| Gemini 3.5 Flash | high | 72.1% | $0.850 | - |
| Gemini 3.5 Flash | minimal | 8.9% | $0.107 | - |
| Claude Sonnet 4.6 | max | 58.3% | $2.72 | - |
| Claude Sonnet 4.6 | high | 60.4% | $2.70 | - |
| Kimi K3 | max | 60.4% | $1.59 | - |
| Kimi K3 | high | 55.0% | $0.947 | - |
| Kimi K3 | low | 12.4% | $0.254 | - |
The best model from each lab
Ordered by the ECI of the most capable model each organization has in the ranking.
OpenAI
Anthropic
Google DeepMind
Moonshot
Meta AI
Alibaba
xAI
Z.ai (Zhipu AI)
DeepSeek
Thinking Machines
Xiaomi Corp
MiniMax
Recent releases, still without a score
Epoch AI registers the model on the day it comes out and computes the capability score later. In that window the model exists, may already be in everyone's hands, and has no ECI yet. It does not take a position in the ranking above, which is ordered by capability, and it is listed here. When the Arena and LiveBench have already measured it, their scores appear; when not, the column stays empty, as in the rest of the page. Today there are 12 models in that situation. The capability column says how many individual Epoch AI tests already have a published result for that model: it is genuine measurement, just not yet consolidated into the index.
| Model | Lab | Release | ELO | LiveBench | Capability | Access |
|---|---|---|---|---|---|---|
| Claude Sonnet 5.5 | Anthropic | Sep 28, 2026 | no match | 77.8 | 7 tests already measured | API access |
| Claude Opus 5.5 | Anthropic | Sep 22, 2026 | 1,572 | 83.2 | 14 tests already measured | API access |
| GPT-6 Luna | OpenAI | Sep 22, 2026 | no match | 72.0 | 12 tests already measured | API access |
| GPT-6 Sol | OpenAI | Sep 22, 2026 | no match | 79.3 | 13 tests already measured | API access |
| Grok 4.7 | xAI | Sep 21, 2026 | 1,476 | 77.4 | 6 tests already measured | API access |
| Step 5 Preview | StepFun | Sep 18, 2026 | no match | no match | 2 tests already measured | API access |
| SWE-2 | Cognition | Sep 10, 2026 | no match | no match | 1 test already measured | Hosted access (no API) |
| Mercury 2.5 | Inception Labs | Sep 8, 2026 | no match | no match | 2 tests already measured | API access |
| Tencent Hy4 preview | Tencent | Aug 28, 2026 | no match | no match | 1 test already measured | Open weights (unrestricted) |
| Qwen3.8 Flash | Alibaba | Aug 26, 2026 | no match | no match | 1 test already measured | API access |
| Grok Build 0.1 | xAI | May 29, 2026 | no match | 67.8 | awaiting measurement | API access |
| GPT-5.2 Codex | OpenAI | Dec 18, 2025 | no match | 74.0 | awaiting measurement | API access |
ROO3 does not measure any model
This page has no lab and runs no benchmark. It gathers, translates and credits two public measurements made by people who have the infrastructure for it. Saying otherwise would be the first lie, and after it no number here would be worth anything.
We order by ECI, not by an average
ECI, ELO and LiveBench are on different scales: the ECI sits around 150, the ELO around 1400 and LiveBench goes from 0 to 100. Adding them into a "score of our own" would produce a nice-looking number that means nothing. The ranking follows the ECI, and the other scores sit alongside for you to compare with your own eyes.
An empty column is an answer, not a failure
The sources write the same model in different ways. Matching by text similarity is exactly how an aggregator starts lying without throwing an error: all it takes is pasting a "Flash" ELO onto the "Pro" with the same name. Here the match is exact or it does not exist. Today, 102 of the 258 models have a confirmed match on the Arena and 48 on LiveBench.
One model, its ceiling
The same model shows up in the sources several times, once per reasoning-effort setting. Instead of repeating the top of the ranking with the same name five times, we keep the best result of each family in each column. It is that model's real ceiling.
A source being down does not empty a column
When one of the sources does not answer, the sync keeps the previous day number and logs the failure. The opposite, silently clearing the column, is the kind of defect nobody notices: the page stays up, looking fine, with one piece of data missing.
Only sources that allow redistribution
There are more famous rankings than these. They were left out because their terms of use forbid republishing the data on a third-party commercial site. We prefer sources that allow it in writing to another one that would bring more traffic and a legal problem.
Six times a day, automatic
A scheduled job downloads the three sources, checks that each license is still the same (CC BY for Epoch AI and the Arena, CC BY-SA for LiveBench), keeps the fingerprint of the day's file and saves. If a license changes, syncing that source stops on its own and nothing new from it is saved.
Whose data this is
Epoch AI and the Arena are under Creative Commons Attribution 4.0, and LiveBench under Attribution-ShareAlike 4.0. All three allow commercial use and redistribution as long as they are credited. The LiveBench numbers stay under the same CC BY-SA. The credit below is generated from the update's own record, not written by hand, so it does not get lost in a future edit.
Epoch AI, "AI Benchmarking Hub". Published at epoch.ai. Data under the CC BY 4.0 licence, reordered by roo3.co.
The Arena (LMArena) leaderboard, from the public lmarena-ai/leaderboard-dataset on Hugging Face, under the CC BY 4.0 licence. We are not affiliated with the Arena.
LiveBench, by the LiveBench team, published at livebench.ai under the CC BY-SA 4.0 license. The overall average is recomputed by roo3.co with the same formula as the website (average of the categories), and those numbers stay under CC BY-SA 4.0. We are not affiliated with LiveBench.
ROO3 is not affiliated with Epoch AI, the Arena (LMArena) or LiveBench. The data was compiled, ordered and translated into Portuguese; the values were not changed, and the LiveBench overall average is recomputed with the same formula as its website. Trademarks belong to their respective owners.
What these numbers mean
Which AI is best today?
It depends on what you mean by best. For capability measured on hard tests, the top of the ECI. For the preference of ordinary people in everyday conversations, the top of the ELO. LiveBench is a third reading, with refreshed questions to reduce the chance that the model has already seen the test. The rankings do not agree, which is why this page shows the columns side by side instead of inventing a single score.
What is the ECI?
Epoch Capabilities Index, from Epoch AI. A composite index that combines a model performance across several hard benchmarks into a single scale, which allows comparing models that never took exactly the same tests. The higher, the better.
What is the Arena ELO?
It is the Arena human preference score. People send the same question to two models without knowing which they are, pick the better answer, and the system calculates a score on the same rating scheme as chess. It measures what people like to receive, which is not the same as raw capability.
What is LiveBench?
It is a benchmark with 23 objective tasks in 7 categories (reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following), with questions refreshed every six months to reduce the chance that the model has already seen the test. The score goes from 0 to 100 and is the average of the categories, the same formula the LiveBench website uses. It usually publishes a release's score before the Epoch AI index.
Why do some models have an empty ELO column?
Because that model was not found in the Arena with certainty. Matching names by similarity is how a ranking starts lying without throwing any error. When there is no exact match, the column stays empty. We prefer the gap to the wrong number.
I cannot find a model that just shipped. Where is it?
If it came out recently, it shows up in the release card at the top of the page and in the "Recent releases, not scored yet" list, right below the table. Epoch AI registers the model on release day and computes the capability score later, so there is a window when the model is public but has no ECI yet. In that window it does not take a position in the ranking, which is ordered by capability, but it appears with its LiveBench and Arena scores when they already exist, which usually happens before the ECI. We do not estimate the missing score.
Do open weight models appear here?
They do, and you can filter for them alone in the bar above the table. The Access column shows whether the model is served through an API, whether the weights are open or whether access is restricted, following Epoch AI own classification.
Can I use this data in my own work?
The original data comes from the three sources credited below: Epoch AI and Arena under Creative Commons Attribution 4.0, and LiveBench under Attribution-ShareAlike 4.0. You can use it, commercially too, as long as you credit the original sources the same way we do here; anything taken from LiveBench stays under the same CC BY-SA. The credit is not a courtesy: it is the condition of the license.
How often does it update?
Six times a day, automatically. The date and time of the last update are at the top of the page and in the table footer. If one of the sources is down, the page keeps the numbers from the previous update instead of blanking the column.
Knowing which AI is best is easy. Using it in your business is the hard part.
We roll AI out in real companies: customer service, content, process. With numbers at the end of the month, not with promises.
Talk about AI in my company