Leaderboard
Measured, dated, sortable
14 models · 1 tracked metric · latest measurement . Click a column to sort. Dashed values are vendor-published and not yet reproduced in our harness.
| Claude Opus 5Anthropic | 63 | $5 | $25 | 1M |
| Claude Fable 5Anthropic | 62 | $10 | $50 | 1M |
| GPT-5.6 SolOpenAI | 61 | $5 | $30 | 1M |
| Kimi K3Moonshot AI | 60 | $3 | $15 | 1M |
| Qwen3.8 MaxAlibaba | 58 | $2 | $6 | 1M |
| GPT-5.6 TerraOpenAI | 57 | $2 | $12 | 1M |
| Muse Spark 1.2Meta | 57 | $1.25 | $4.25 | 1M |
| Grok 4.5xAI (SpaceX) | 56 | $2 | $6 | 512K |
| Claude Sonnet 5Anthropic | 55 | $2 | $10 | 1M |
| GLM-5.2Z AI | 53 | $1.4 | $4.4 | 1M |
| DeepSeek V4 FlashDeepSeek | 52 | $0.14 | $0.28 | 1M |
| Gemini 3.6 FlashGoogle | 52 | $1.5 | $7.5 | 1M |
| GPT-5.6 LunaOpenAI | 52 | $0.2 | $1.2 | 1M |
| Gemini 3.1 Pro PreviewGoogle | 48 | $2 | $12 | 1M |
Methodology note
The primary metric is the Artificial Analysis Intelligence Index — an independent third-party composite of ten evaluations — recorded here with the date we verified it. Prices are the rates published on each model's Artificial Analysis page as of the same date. Columns from our own measurement harness (SWE-bench Verified, Terminal-Bench) will join as those runs land; the harness protocol is in How we benchmark models on llm.blog.
