News · tag
benchmarks
2 posts
August refresh: all 12 tracked models re-scored on SWE-bench VerifiedDraft
Claude Opus 4.5 holds the top slot; Gemini 3 Flash posts the largest gain of the cycle. Full deltas on the leaderboard.
Terminal-Bench 2.0 joins the llm.blog tracked benchmark setDraft
Six models scored on real terminal-environment tasks in the July cycle. Why we picked it, and what it measures that SWE-bench doesn't.