Skip to content

News · leaderboard · methodology · pricing

llm.blog dataset refresh: 14 current models, AA Intelligence Index adopted

llm.blog rebuilt its model dataset against Artificial Analysis on August 11, 2026: 14 current models with verified pricing and the AA Intelligence Index as primary metric. Claude Opus 5 leads at 63, Claude Fable 5 follows at 62, Kimi K3 tops open weights at 60, and DeepSeek V4 Flash holds the price floor at $0.14 per million input tokens.

By Kushagra Sikka1 min read

We rebuilt the llm.blog model dataset on August 11, 2026, against Artificial Analysis — replacing the prior tracked set, whose specifications and prices no longer reflected the market. Fourteen current models now carry verified per-MTok pricing, context windows, and the AA Intelligence Index as the leaderboard’s primary metric, each value dated to the day we verified it.

The state of the frontier, measured

The top of the index as of this refresh:

Model Maker Index $/MTok in · out
Claude Opus 5 Anthropic 63 $5 · $25
Claude Fable 5 Anthropic 62 $10 · $50
Kimi K3 Moonshot 60 $3 · $15
GPT-5.6 Sol OpenAI 61 $5 · $30
Qwen3.8 Max Alibaba 58 $2 · $6

Three facts stand out. Claude Opus 5 leads the index at 63 while undercutting its own premium sibling — Fable 5 costs twice as much and measures a point lower. Kimi K3, a 2.8-trillion-parameter open-weights MoE, sits three points off the closed frontier — the narrowest open-closed gap yet measured. And DeepSeek V4 Flash at $0.14/$0.28 per million tokens delivers an index score of 52 — matching GPT-5.6 Luna and Gemini 3.6 Flash — at prices Artificial Analysis rates the lowest of any well-known model.

What changed methodologically

The AA Intelligence Index — an independent composite of ten evaluations — is now the leaderboard’s primary column, recorded with the date we verified it rather than presented as our own measurement. Columns from our own harness (SWE-bench Verified, Terminal-Bench) will join as those runs land; the harness protocol is unchanged in the methodology. Pricing now cites each model’s Artificial Analysis page as of the verification date.

The full table, sortable with per-cell dates, is on the leaderboard. All fifteen head-to-head comparisons were rewritten against the new data the same day.