News
Releases, prices, measurements
Short, timestamped posts on what changed. Every fact carries a date; every number carries a source.
llm.blog dataset refresh: 14 current models, AA Intelligence Index adopted
The tracked model set is rebuilt against Artificial Analysis data as of August 11, 2026: Claude Opus 5 leads at 63, Kimi K3 tops open weights, and DeepSeek V4 Flash sets the price floor.
Meta's Muse Spark 1.2 and Alibaba's Qwen3.8 Max land in the frontier top ten
Two releases in three days reshuffle the mid-frontier: Meta goes proprietary with full multimodal input, and Alibaba undercuts every flagship at $2/$6 per million tokens.
Google cuts Gemini 3 Flash input pricing 17% to $0.25 per million tokensDraft
The second Flash-tier price cut this year puts Google level with GPT-5 mini on input cost — with a 1M-token context window.
Anthropic promotes Claude Opus 4.5 snapshot 20260805 to API defaultDraft
The new default snapshot tightens tool-use formatting; Anthropic reports no benchmark movement. Pinned versions keep working.
August refresh: all 12 tracked models re-scored on SWE-bench VerifiedDraft
Claude Opus 4.5 holds the top slot; Gemini 3 Flash posts the largest gain of the cycle. Full deltas on the leaderboard.
Kimi K2 Thinking adds native tool-call streaming in the Moonshot APIDraft
Tool calls now stream token-by-token instead of arriving as completed blocks — closing a latency gap with closed-model APIs.
OpenAI halves GPT-5.1 cached-input pricing to $0.125 per million tokensDraft
Cached input now costs 10% of the base rate — a direct subsidy for long-system-prompt agents and high-frequency RAG.
Alibaba ships Qwen3-Max 2026-07 update with improved long-context recallDraft
The mid-cycle update targets retrieval degradation past 128K tokens; pricing and API surface are unchanged.
Terminal-Bench 2.0 joins the llm.blog tracked benchmark setDraft
Six models scored on real terminal-environment tasks in the July cycle. Why we picked it, and what it measures that SWE-bench doesn't.
DeepSeek extends off-peak API discount window to 12 hours dailyDraft
Half-price inference now covers the full UTC night — and batch schedulers are already migrating.
Claude Sonnet 4.5's 1M-token context window graduates from betaDraft
Long-context pricing kicks in above 200K tokens. The beta header is gone; the premium tier is now the story.
