News · benchmarks · leaderboard · swe-bench
August refresh: all 12 tracked models re-scored on SWE-bench Verified
Draft — pending editorial review
llm.blog re-measured all 12 tracked models on SWE-bench Verified during the first week of August 2026. Claude Opus 4.5 leads at 80.9%. Gemini 3 Flash posted the cycle's largest gain. Every score on the leaderboard now carries an August 5, 2026 measurement date and harness configuration notes.
We completed the August SWE-bench Verified refresh on August 5, 2026, re-running all 12 tracked models through the harness described in our methodology. Every leaderboard cell now carries this cycle’s measurement date.
Results snapshot
The ordering at the top is unchanged from July: Claude Opus 4.5 leads at 80.9%, with Claude Sonnet 4.5 second at 77.2%. Movement concentrated in the middle of the table:
- Gemini 3 Flash posted the cycle’s largest gain, consistent with Google’s July serving-stack update.
- Kimi K2 Thinking and GPT-5.1 were stable within run-to-run variance (±0.4 points across three seeds).
- No model regressed beyond variance bounds this cycle.
Full per-model numbers, with dates and configuration, are on the leaderboard.
Measurement notes
Runs used vendor-default reasoning settings, a 128K context ceiling for parity, and the agentless-lite scaffold documented in the methodology post. Three seeds per model; we report the median. Scores for models whose vendors publish higher numbers under custom scaffolds are flagged unverified on model pages until we can reproduce them.
One harness change this cycle: we upgraded the container image to match SWE-bench’s July revision, which resolved two flaky test environments in the django subset. The change affected all models identically and moved no median by more than 0.2 points.
September’s refresh is scheduled for the first week of the month; subscribe to the changelog via RSS to get refresh announcements as they land.