Skip to content

Leaderboard

Measured, dated, sortable

14 models · 1 tracked metric · latest measurement . Click a column to sort. Dashed values are vendor-published and not yet reproduced in our harness.

Claude Opus 5Anthropic63$5$251M
Claude Fable 5Anthropic62$10$501M
GPT-5.6 SolOpenAI61$5$301M
Kimi K3Moonshot AI60$3$151M
Qwen3.8 MaxAlibaba58$2$61M
GPT-5.6 TerraOpenAI57$2$121M
Muse Spark 1.2Meta57$1.25$4.251M
Grok 4.5xAI (SpaceX)56$2$6512K
Claude Sonnet 5Anthropic55$2$101M
GLM-5.2Z AI53$1.4$4.41M
DeepSeek V4 FlashDeepSeek52$0.14$0.281M
Gemini 3.6 FlashGoogle52$1.5$7.51M
GPT-5.6 LunaOpenAI52$0.2$1.21M
Gemini 3.1 Pro PreviewGoogle48$2$121M

Methodology note

The primary metric is the Artificial Analysis Intelligence Index — an independent third-party composite of ten evaluations — recorded here with the date we verified it. Prices are the rates published on each model's Artificial Analysis page as of the same date. Columns from our own measurement harness (SWE-bench Verified, Terminal-Bench) will join as those runs land; the harness protocol is in How we benchmark models on llm.blog.