News · tag
methodology
2 posts
llm.blog dataset refresh: 14 current models, AA Intelligence Index adopted
The tracked model set is rebuilt against Artificial Analysis data as of August 11, 2026: Claude Opus 5 leads at 63, Kimi K3 tops open weights, and DeepSeek V4 Flash sets the price floor.
Terminal-Bench 2.0 joins the llm.blog tracked benchmark setDraft
Six models scored on real terminal-environment tasks in the July cycle. Why we picked it, and what it measures that SWE-bench doesn't.