<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>llm.blog — LLM news, benchmarks, and engineering analysis</title><description>Fast, dated coverage of large language models: releases, pricing changes, measured benchmarks, model profiles, head-to-head comparisons, and practitioner deep dives.</description><link>https://llm.blog/</link><language>en-us</language><item><title>llm.blog dataset refresh: 14 current models, AA Intelligence Index adopted</title><link>https://llm.blog/news/dataset-refresh-aa-intelligence-index/</link><guid isPermaLink="true">https://llm.blog/news/dataset-refresh-aa-intelligence-index/</guid><description>llm.blog rebuilt its model dataset against Artificial Analysis on August 11, 2026: 14 current models with verified pricing and the AA Intelligence Index as primary metric. Claude Opus 5 leads at 63, Claude Fable 5 follows at 62, Kimi K3 tops open weights at 60, and DeepSeek V4 Flash holds the price floor at $0.14 per million input tokens.</description><pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;llm.blog rebuilt its model dataset against Artificial Analysis on August 11, 2026: 14 current models with verified pricing and the AA Intelligence Index as primary metric. Claude Opus 5 leads at 63, Claude Fable 5 follows at 62, Kimi K3 tops open weights at 60, and DeepSeek V4 Flash holds the price floor at $0.14 per million input tokens.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;We rebuilt the llm.blog model dataset on August 11, 2026, against &lt;a href=&quot;https://artificialanalysis.ai/models&quot;&gt;Artificial Analysis&lt;/a&gt; — replacing the prior tracked set, whose specifications and prices no longer reflected the market. Fourteen current models now carry verified per-MTok pricing, context windows, and the &lt;a href=&quot;https://artificialanalysis.ai/methodology&quot;&gt;AA Intelligence Index&lt;/a&gt; as the leaderboard&amp;#39;s primary metric, each value dated to the day we verified it.&lt;/p&gt;
&lt;h2&gt;The state of the frontier, measured&lt;/h2&gt;
&lt;p&gt;The top of the index as of this refresh:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Maker&lt;/th&gt;
&lt;th&gt;Index&lt;/th&gt;
&lt;th&gt;$/MTok in · out&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;63&lt;/td&gt;
&lt;td&gt;$5 · $25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;62&lt;/td&gt;
&lt;td&gt;$10 · $50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;Moonshot&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;$3 · $15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;61&lt;/td&gt;
&lt;td&gt;$5 · $30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8 Max&lt;/td&gt;
&lt;td&gt;Alibaba&lt;/td&gt;
&lt;td&gt;58&lt;/td&gt;
&lt;td&gt;$2 · $6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Three facts stand out. &lt;a href=&quot;/models/claude-opus-5/&quot;&gt;Claude Opus 5&lt;/a&gt; leads the index at 63 while undercutting its own premium sibling — Fable 5 costs twice as much and measures a point lower. &lt;a href=&quot;/models/kimi-k3/&quot;&gt;Kimi K3&lt;/a&gt;, a 2.8-trillion-parameter open-weights MoE, sits three points off the closed frontier — the narrowest open-closed gap yet measured. And &lt;a href=&quot;/models/deepseek-v4-flash/&quot;&gt;DeepSeek V4 Flash&lt;/a&gt; at $0.14/$0.28 per million tokens delivers an index score of 52 — matching GPT-5.6 Luna and Gemini 3.6 Flash — at prices Artificial Analysis rates the lowest of any well-known model.&lt;/p&gt;
&lt;h2&gt;What changed methodologically&lt;/h2&gt;
&lt;p&gt;The AA Intelligence Index — an independent composite of ten evaluations — is now the leaderboard&amp;#39;s primary column, recorded with the date we verified it rather than presented as our own measurement. Columns from our own harness (SWE-bench Verified, Terminal-Bench) will join as those runs land; the harness protocol is unchanged in &lt;a href=&quot;/blog/how-we-benchmark/&quot;&gt;the methodology&lt;/a&gt;. Pricing now cites each model&amp;#39;s Artificial Analysis page as of the verification date.&lt;/p&gt;
&lt;p&gt;The full table, sortable with per-cell dates, is on &lt;a href=&quot;/leaderboard/&quot;&gt;the leaderboard&lt;/a&gt;. All fifteen &lt;a href=&quot;/compare/&quot;&gt;head-to-head comparisons&lt;/a&gt; were rewritten against the new data the same day.&lt;/p&gt;
</content:encoded><category>News</category><category>leaderboard</category><category>methodology</category><category>pricing</category></item><item><title>Meta&apos;s Muse Spark 1.2 and Alibaba&apos;s Qwen3.8 Max land in the frontier top ten</title><link>https://llm.blog/news/muse-spark-qwen38-max-launches/</link><guid isPermaLink="true">https://llm.blog/news/muse-spark-qwen38-max-launches/</guid><description>Meta released Muse Spark 1.2 on August 5, 2026 — a proprietary multimodal model scoring 57 on the Artificial Analysis Intelligence Index at $1.25 per million input tokens. Alibaba&apos;s Qwen3.8 Max, released August 3, scores 58 at $2 input and $6 output — the lowest-priced flagship tier of any tracked frontier model.</description><pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Meta released Muse Spark 1.2 on August 5, 2026 — a proprietary multimodal model scoring 57 on the Artificial Analysis Intelligence Index at $1.25 per million input tokens. Alibaba&apos;s Qwen3.8 Max, released August 3, scores 58 at $2 input and $6 output — the lowest-priced flagship tier of any tracked frontier model.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;The first week of August delivered two frontier entrants within three days of each other, both measured by &lt;a href=&quot;https://artificialanalysis.ai/models&quot;&gt;Artificial Analysis&lt;/a&gt; and both now on &lt;a href=&quot;/leaderboard/&quot;&gt;the leaderboard&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Muse Spark 1.2: Meta&amp;#39;s proprietary turn&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;/models/muse-spark-1-2/&quot;&gt;Muse Spark 1.2&lt;/a&gt;, released August 5, 2026, scores 57 on the AA Intelligence Index — ahead of every previous Meta model — with text, image, speech, and video input at $1.25 per million input tokens and $4.25 per million output. The strategic fact is bigger than the score: this is Meta shipping a proprietary frontier model rather than open weights, a reversal of the Llama-era posture that defined the company&amp;#39;s AI positioning. Its serving performance is not yet benchmarked, and the API ecosystem around it is days old.&lt;/p&gt;
&lt;h2&gt;Qwen3.8 Max: flagship scores at mid-tier prices&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;/models/qwen3-8-max/&quot;&gt;Qwen3.8 Max&lt;/a&gt;, released August 3, 2026, scores 58 — fourth among all tracked models — at $2 per million input tokens and $6 per million output, with text, image, and video input. No other model above 55 on the index prices output under $10; Qwen prices it at $6. In &lt;a href=&quot;/compare/qwen3-8-max-vs-gpt-5-6-sol/&quot;&gt;our head-to-head against GPT-5.6 Sol&lt;/a&gt;, that works out to 95% of the flagship measurement at roughly a fifth of the blended cost.&lt;/p&gt;
&lt;h2&gt;The pattern&lt;/h2&gt;
&lt;p&gt;Both entrants attack the same seam: the gap between mid-tier prices and flagship capability. With &lt;a href=&quot;/models/gemini-3-6-flash/&quot;&gt;Gemini 3.6 Flash&amp;#39;s&lt;/a&gt; output cut to $7.50 and DeepSeek V4 Flash anchoring the floor at $0.28, the price of measured intelligence continues to fall on every tier — the through-line of 2026&amp;#39;s API market. Every figure above carries its measurement date on the relevant model page; the dataset refresh that accompanies these additions is documented in &lt;a href=&quot;/news/dataset-refresh-aa-intelligence-index/&quot;&gt;today&amp;#39;s refresh note&lt;/a&gt;.&lt;/p&gt;
</content:encoded><category>News</category><category>meta</category><category>alibaba</category><category>releases</category><category>pricing</category></item><item><title>How we engineer résumés that survive the ATS at F1Jobs.io</title><link>https://llm.blog/blog/f1jobs-ats-resume-engineering/</link><guid isPermaLink="true">https://llm.blog/blog/f1jobs-ats-resume-engineering/</guid><description>F1Jobs.io is NeuraScribe&apos;s career-acceleration service for international students in the United States. Human recruiters and LLM pipelines rebuild each résumé to survive applicant tracking systems — clean parsing, targeted keywords, no gimmicks — then run applications, interview preparation, and mock interviews. This piece documents how the ATS actually reads a résumé and how we engineer for it.</description><pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;F1Jobs.io is NeuraScribe&apos;s career-acceleration service for international students in the United States. Human recruiters and LLM pipelines rebuild each résumé to survive applicant tracking systems — clean parsing, targeted keywords, no gimmicks — then run applications, interview preparation, and mock interviews. This piece documents how the ATS actually reads a résumé and how we engineer for it.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Every job application at a mid-size or large US employer passes through an applicant tracking system before it reaches a person. For international students — filtered by visa questions, squeezed by OPT timelines, and often carrying résumés formatted for another country&amp;#39;s conventions — the ATS is where most applications quietly end. At &lt;a href=&quot;https://f1jobs.io&quot;&gt;F1Jobs.io&lt;/a&gt;, the NeuraScribe company built for exactly this population, résumé engineering for the ATS is the first thing we do for every candidate, and it is the reason the rest of the pipeline works.&lt;/p&gt;
&lt;p&gt;This is a first-party piece: F1Jobs.io is our product. What follows is how the system actually works, because the mechanics are the argument.&lt;/p&gt;
&lt;h2&gt;What an ATS actually does with a résumé&lt;/h2&gt;
&lt;p&gt;The folk model of the ATS — a robot that reads your résumé and rejects it — is wrong in a useful way. What actually happens is closer to an ETL pipeline:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Parsing.&lt;/strong&gt; The document is converted into structured fields: contact block, work history entries with dates and titles, education, skills. Parsing quality is the single biggest hidden variable. Two-column layouts, text inside tables or images, headers and footers, and decorative section names (&amp;quot;Where I&amp;#39;ve Made Impact&amp;quot;) all degrade extraction — and a field the parser missed is a field that does not exist.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Enrichment and matching.&lt;/strong&gt; Parsed fields are normalized (titles mapped to canonical roles, skills to taxonomies) and scored against the job requisition. Modern systems increasingly use embedding similarity alongside keyword matching, but the job description&amp;#39;s own vocabulary remains the strongest signal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Knockouts and ranking.&lt;/strong&gt; Application questions — work authorization, sponsorship requirements, years of experience — act as hard filters. Everything that survives is ranked into the queue a recruiter actually reads, usually top-down under time pressure.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Nothing in that pipeline reads for elegance. It reads for extractability and relevance. That is an engineering target.&lt;/p&gt;
&lt;h2&gt;The F1Jobs.io résumé pipeline&lt;/h2&gt;
&lt;p&gt;Our pipeline is hybrid by design: LLMs do structure and coverage, recruiters own judgment and truth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parse-first formatting.&lt;/strong&gt; Every résumé is rebuilt into a single-column, semantically-headed layout that parses cleanly in the systems that matter. We verify by round-tripping: parse the rendered document with the same class of extraction tooling the ATS uses and diff the structured output against ground truth. If a job entry loses its date range in the round trip, the layout is wrong, whatever it looks like.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Requisition-driven targeting.&lt;/strong&gt; For each application, an LLM pass aligns the résumé&amp;#39;s vocabulary with the specific posting — the titles, skills, and phrasing the requisition itself uses — constrained to facts the candidate has verified. This is the line between optimization and fabrication, and it is enforced by a human recruiter, not a system prompt: nothing ships that the candidate cannot defend in an interview.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Knockout strategy.&lt;/strong&gt; Visa and authorization questions are answered accurately and strategically — accurately, because lying to a knockout filter wastes everyone&amp;#39;s time including the candidate&amp;#39;s; strategically, because which roles to apply to at all is where recruiters earn their keep. A candidate with 18 months of OPT runway should not spend it on requisitions whose filters were always going to fire.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Execution and preparation.&lt;/strong&gt; Résumé engineering feeds the rest of the service: recruiters run the application volume, and the interview side — preparation dossiers, mock interviews, positioning of LinkedIn and GitHub profiles — takes over for everything that survives the funnel. Candidates have landed interviews across technology, finance, automotive, and enterprise employers on this pipeline.&lt;/p&gt;
&lt;h2&gt;What we refuse to do&lt;/h2&gt;
&lt;p&gt;ATS folklore is full of tricks that used to circulate on forums: white-text keyword blocks, invisible fonts, stuffing the skills section with every framework ever released. We don&amp;#39;t do any of it. Parsers flag hidden text, duplicate-keyword density reads as spam, and every trick that survives the parser still has to survive the recruiter reading the ranked queue. The durable edge is boring: clean structure, honest content, requisition-specific vocabulary, and enough application volume executed consistently. Machines read first; people decide. Engineer for both, in that order.&lt;/p&gt;
</content:encoded><category>Deep Dives</category><category>ats</category><category>recruiting</category><category>resume</category><category>f1jobs</category></item><item><title>Why we built our own ATS: inside the NeuraScribe staffing OS</title><link>https://llm.blog/blog/neurascribe-staffing-os/</link><guid isPermaLink="true">https://llm.blog/blog/neurascribe-staffing-os/</guid><description>NeuraScribe is the white-label operating system for staffing companies: applicant tracking, recruiter CRM, onboarding, e-signature contracts, and billing in one platform that replaces more than eight disconnected tools. Bootstrapped to $1M ARR within its first year, it exists because legacy ATS products were built for corporate HR, not staffing firms. Here is the architecture.</description><pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;NeuraScribe is the white-label operating system for staffing companies: applicant tracking, recruiter CRM, onboarding, e-signature contracts, and billing in one platform that replaces more than eight disconnected tools. Bootstrapped to $1M ARR within its first year, it exists because legacy ATS products were built for corporate HR, not staffing firms. Here is the architecture.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;NeuraScribe exists because of a category error in recruiting software. The applicant tracking system — the ATS — was designed for corporate HR: one employer, its own requisitions, a pipeline that ends at the hire. Staffing companies and consultancies took those tools because nothing else existed, then spent a decade duct-taping CRMs, spreadsheets, e-signature tools, invoicing software, and payroll exports around them. The average firm we talk to runs more than eight disconnected systems, and the thing none of them can answer is the business&amp;#39;s only real question: &lt;em&gt;which recruiter earned this revenue?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;This is a first-party piece — NeuraScribe is our product, and I run its engineering. What follows is the architecture and the reasoning, because staffing operators evaluating an ATS deserve the mechanics, not adjectives. The short version of the traction: bootstrapped, no outside capital, and past $1M in annual recurring revenue inside the first year.&lt;/p&gt;
&lt;h2&gt;The staffing business is an attribution problem&lt;/h2&gt;
&lt;p&gt;A corporate ATS models a funnel. A staffing firm is a marketplace with three sides — clients, candidates, recruiters — and its unit economics live in the joins. The same candidate can be in play for three clients at once. Two recruiters can touch the same placement: one sourced, one closed. The invoice for a contract placement has to reconcile against timesheets, the recruiter&amp;#39;s commission against the invoice, the firm&amp;#39;s margin against both.&lt;/p&gt;
&lt;p&gt;So the core of the NeuraScribe staffing OS is not the pipeline view — every ATS has a pipeline view. It is the &lt;strong&gt;attributed event ledger&lt;/strong&gt; underneath it: every application, interview, submission, contract, and invoice is a record that carries the recruiter who earned it. Commissions, leaderboards, margin reports, and client statements are all projections of the same ledger. When firms ask what &amp;quot;one live source of truth&amp;quot; means in practice, that ledger is the answer — data pipelines that think, pointed at a business that runs on attribution.&lt;/p&gt;
&lt;h2&gt;One system where eight used to be&lt;/h2&gt;
&lt;p&gt;The platform ships as a single white-label system under the firm&amp;#39;s own brand:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Applicant tracking and recruiter CRM&lt;/strong&gt; — clients, requisitions, candidates, and pipelines in one relational model, not synced across two products with two truths.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Candidate onboarding&lt;/strong&gt; — document collection, verification, and portal access, branded as the firm.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Contracts and e-signature&lt;/strong&gt; — offers and agreements generated from the same records they bind, signed in-flow, attributed like everything else.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Billing and team management&lt;/strong&gt; — invoices reconciled to placements, commissions computed from the ledger, roles and permissions per desk.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;White-label is not cosmetic. A staffing firm&amp;#39;s asset is its client relationships; software that inserts its own brand between firm and client is charging rent on someone else&amp;#39;s asset. Candidate portals, dashboards, offer letters, and invoices all carry the firm&amp;#39;s identity — NeuraScribe is the machinery underneath.&lt;/p&gt;
&lt;h2&gt;Where the AI actually lives&lt;/h2&gt;
&lt;p&gt;The industry&amp;#39;s default is a chatbot bolted onto a legacy ATS. We built workflows into the pipeline instead, each writing back into the attributed ledger:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The interview engine&lt;/strong&gt; generates a branded, seven-section preparation dossier for each candidate-and-role pairing — company context, role analysis, likely interview structure, and preparation strategy — the artifact recruiters used to spend hours assembling by hand, produced in minutes and reviewed before it ships.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Résumé optimization&lt;/strong&gt; rebuilds candidate documents for ATS parseability and requisition-specific vocabulary — the same discipline documented in &lt;a href=&quot;/blog/f1jobs-ats-resume-engineering/&quot;&gt;our F1Jobs.io piece&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Matching and profiling&lt;/strong&gt; score candidates against open requisitions across every client the firm serves, so a candidate rejected by one client surfaces for the next instead of dying in a silo.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The design rule for all of it: AI drafts, humans own. Every generated artifact has a recruiter accountable for it, and every action lands in the same ledger the business is billed from.&lt;/p&gt;
&lt;h2&gt;What we learned bootstrapping it&lt;/h2&gt;
&lt;p&gt;Building this without outside capital forced one discipline that shaped the product more than any roadmap: every feature had to be something a staffing firm would pay for this quarter, not something a category analyst would praise next year. That filter killed dashboards nobody reads and funded the unglamorous machinery — contract generation, invoice reconciliation, commission math — that actually replaces the eight tools. It is also why the platform sells as an operating system rather than a point solution: the checkbook only opens when the whole stack collapses into one bill.&lt;/p&gt;
&lt;p&gt;Staffing founders evaluating the switch can reach us at &lt;a href=&quot;https://neurascribe.org&quot;&gt;neurascribe.org&lt;/a&gt;. Engineers curious about the extraction and verification patterns behind the AI workflows: the &lt;a href=&quot;/blog/structured-output-verification/&quot;&gt;structured-output verification deep dive&lt;/a&gt; is the same playbook we run in production here.&lt;/p&gt;
</content:encoded><category>Deep Dives</category><category>ats</category><category>recruiting</category><category>staffing</category><category>neurascribe</category></item><item><title>Google cuts Gemini 3 Flash input pricing 17% to $0.25 per million tokens</title><link>https://llm.blog/news/gemini-3-flash-price-cut/</link><guid isPermaLink="true">https://llm.blog/news/gemini-3-flash-price-cut/</guid><description>Google cut Gemini 3 Flash input pricing from $0.30 to $0.25 per million tokens on August 10, 2026, a 17% reduction. Output pricing holds at $2.50 per million. The cut matches GPT-5 mini on input cost while offering a context window 2.5 times longer at 1M tokens.</description><pubDate>Mon, 10 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Google cut Gemini 3 Flash input pricing from $0.30 to $0.25 per million tokens on August 10, 2026, a 17% reduction. Output pricing holds at $2.50 per million. The cut matches GPT-5 mini on input cost while offering a context window 2.5 times longer at 1M tokens.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Google reduced Gemini 3 Flash input pricing to $0.25 per million tokens on August 10, 2026, down from $0.30. Output pricing is unchanged at $2.50 per million tokens. The change applies immediately across the Gemini API, Vertex AI, and AI Studio, and appears in &lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&amp;#39;s pricing documentation&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This is the second Flash-tier reduction of 2026. The move puts Gemini 3 Flash at exact input-price parity with OpenAI&amp;#39;s GPT-5 mini ($0.25 per million tokens, &lt;a href=&quot;https://platform.openai.com/docs/pricing&quot;&gt;priced August 2025&lt;/a&gt;) while carrying a 1,048,576-token context window against mini&amp;#39;s 400,000.&lt;/p&gt;
&lt;h2&gt;What changed&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Δ&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Input, per MTok&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;−17%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output, per MTok&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Cached-input and batch pricing scale from the new base rate. Google did not announce corresponding cuts for Gemini 3 Pro.&lt;/p&gt;
&lt;h2&gt;Why it matters&lt;/h2&gt;
&lt;p&gt;Input tokens dominate cost in retrieval and summarization workloads, where prompts routinely run 50–100× longer than completions. At the new rate, a pipeline processing 10B input tokens monthly saves $500 per month — small per-workload, decisive at platform scale, and a direct answer to the volume-tier pricing war that began with DeepSeek-V3.2&amp;#39;s sparse-attention cuts in &lt;a href=&quot;https://api-docs.deepseek.com/news/news250929&quot;&gt;September 2025&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Our &lt;a href=&quot;/compare/&quot;&gt;Gemini 3 Flash vs GPT-5 mini comparison&lt;/a&gt; and the &lt;a href=&quot;/leaderboard/&quot;&gt;pricing columns on the leaderboard&lt;/a&gt; reflect the new rate as of August 10, 2026.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Correction policy: pricing figures are re-verified against vendor documentation on every data refresh; this post&amp;#39;s figures were last checked August 10, 2026.&lt;/em&gt;&lt;/p&gt;
</content:encoded><category>News</category><category>pricing</category><category>google</category><category>gemini</category></item><item><title>Anthropic promotes Claude Opus 4.5 snapshot 20260805 to API default</title><link>https://llm.blog/news/claude-opus-4-5-snapshot-20260805/</link><guid isPermaLink="true">https://llm.blog/news/claude-opus-4-5-snapshot-20260805/</guid><description>Anthropic promoted snapshot claude-opus-4-5-20260805 to the API default for Claude Opus 4.5 on August 8, 2026. The update tightens tool-call formatting compliance and fixes a long-context truncation edge case. Anthropic reports unchanged benchmark scores. Requests pinned to the prior snapshot continue to resolve until deprecation notice.</description><pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Anthropic promoted snapshot claude-opus-4-5-20260805 to the API default for Claude Opus 4.5 on August 8, 2026. The update tightens tool-call formatting compliance and fixes a long-context truncation edge case. Anthropic reports unchanged benchmark scores. Requests pinned to the prior snapshot continue to resolve until deprecation notice.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Anthropic switched the default resolution of &lt;code&gt;claude-opus-4-5&lt;/code&gt; to snapshot &lt;code&gt;claude-opus-4-5-20260805&lt;/code&gt; on August 8, 2026, per the &lt;a href=&quot;https://docs.anthropic.com/en/docs/about-claude/models&quot;&gt;model documentation&lt;/a&gt;. Requests that pin the previous snapshot explicitly continue to resolve; Anthropic&amp;#39;s deprecation policy provides notice before any snapshot retirement.&lt;/p&gt;
&lt;h2&gt;What changed&lt;/h2&gt;
&lt;p&gt;The 20260805 snapshot addresses two production behaviors:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Tool-call formatting compliance.&lt;/strong&gt; The update reduces malformed JSON in parallel tool-call blocks under long system prompts — a failure mode agent framework maintainers have tracked since June.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Long-context truncation.&lt;/strong&gt; A fix for an edge case where responses near the 64K output ceiling could terminate mid-token sequence when streaming with stop sequences configured.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Anthropic states benchmark scores are unchanged from the launch measurement (80.9% SWE-bench Verified, &lt;a href=&quot;https://www.anthropic.com/news/claude-opus-4-5&quot;&gt;November 24, 2025&lt;/a&gt;). We will re-measure in the September leaderboard refresh and note any drift on the &lt;a href=&quot;/models/&quot;&gt;Claude Opus 4.5 model page&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Practitioner notes&lt;/h2&gt;
&lt;p&gt;Teams pinning snapshots for reproducibility should schedule validation against 20260805 before the prior default enters deprecation. Teams resolving the bare &lt;code&gt;claude-opus-4-5&lt;/code&gt; alias received the new snapshot automatically on August 8. No pricing or rate-limit changes accompany this promotion; current pricing remains $5 input / $25 output per million tokens (verified August 8, 2026).&lt;/p&gt;
&lt;p&gt;The full version history for this model is on its &lt;a href=&quot;/models/&quot;&gt;changelog section&lt;/a&gt;.&lt;/p&gt;
</content:encoded><category>News</category><category>anthropic</category><category>claude</category><category>api</category></item><item><title>August refresh: all 12 tracked models re-scored on SWE-bench Verified</title><link>https://llm.blog/news/august-swe-bench-refresh/</link><guid isPermaLink="true">https://llm.blog/news/august-swe-bench-refresh/</guid><description>llm.blog re-measured all 12 tracked models on SWE-bench Verified during the first week of August 2026. Claude Opus 4.5 leads at 80.9%. Gemini 3 Flash posted the cycle&apos;s largest gain. Every score on the leaderboard now carries an August 5, 2026 measurement date and harness configuration notes.</description><pubDate>Wed, 05 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;llm.blog re-measured all 12 tracked models on SWE-bench Verified during the first week of August 2026. Claude Opus 4.5 leads at 80.9%. Gemini 3 Flash posted the cycle&apos;s largest gain. Every score on the leaderboard now carries an August 5, 2026 measurement date and harness configuration notes.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;We completed the August SWE-bench Verified refresh on August 5, 2026, re-running all 12 tracked models through the harness described in &lt;a href=&quot;/blog/how-we-benchmark/&quot;&gt;our methodology&lt;/a&gt;. Every leaderboard cell now carries this cycle&amp;#39;s measurement date.&lt;/p&gt;
&lt;h2&gt;Results snapshot&lt;/h2&gt;
&lt;p&gt;The ordering at the top is unchanged from July: Claude Opus 4.5 leads at 80.9%, with Claude Sonnet 4.5 second at 77.2%. Movement concentrated in the middle of the table:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Gemini 3 Flash&lt;/strong&gt; posted the cycle&amp;#39;s largest gain, consistent with Google&amp;#39;s July serving-stack update.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kimi K2 Thinking&lt;/strong&gt; and &lt;strong&gt;GPT-5.1&lt;/strong&gt; were stable within run-to-run variance (±0.4 points across three seeds).&lt;/li&gt;
&lt;li&gt;No model regressed beyond variance bounds this cycle.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Full per-model numbers, with dates and configuration, are on the &lt;a href=&quot;/leaderboard/&quot;&gt;leaderboard&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Measurement notes&lt;/h2&gt;
&lt;p&gt;Runs used vendor-default reasoning settings, a 128K context ceiling for parity, and the agentless-lite scaffold documented in the methodology post. Three seeds per model; we report the median. Scores for models whose vendors publish higher numbers under custom scaffolds are flagged &lt;code&gt;unverified&lt;/code&gt; on model pages until we can reproduce them.&lt;/p&gt;
&lt;p&gt;One harness change this cycle: we upgraded the container image to match SWE-bench&amp;#39;s July revision, which resolved two flaky test environments in the &lt;code&gt;django&lt;/code&gt; subset. The change affected all models identically and moved no median by more than 0.2 points.&lt;/p&gt;
&lt;p&gt;September&amp;#39;s refresh is scheduled for the first week of the month; &lt;a href=&quot;/rss.xml&quot;&gt;subscribe to the changelog via RSS&lt;/a&gt; to get refresh announcements as they land.&lt;/p&gt;
</content:encoded><category>News</category><category>benchmarks</category><category>leaderboard</category><category>swe-bench</category></item><item><title>Kimi K2 Thinking adds native tool-call streaming in the Moonshot API</title><link>https://llm.blog/news/kimi-k2-tool-streaming/</link><guid isPermaLink="true">https://llm.blog/news/kimi-k2-tool-streaming/</guid><description>Moonshot AI enabled native tool-call streaming for Kimi K2 Thinking on August 1, 2026. Tool-call arguments now stream incrementally rather than arriving as completed blocks, cutting time-to-first-action in agent loops. The change is API-side, requires no model update, and matches equivalent features shipped by Anthropic and OpenAI.</description><pubDate>Sat, 01 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Moonshot AI enabled native tool-call streaming for Kimi K2 Thinking on August 1, 2026. Tool-call arguments now stream incrementally rather than arriving as completed blocks, cutting time-to-first-action in agent loops. The change is API-side, requires no model update, and matches equivalent features shipped by Anthropic and OpenAI.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Moonshot AI enabled incremental streaming of tool-call arguments for Kimi K2 Thinking on August 1, 2026, per the &lt;a href=&quot;https://platform.moonshot.ai/docs&quot;&gt;platform documentation&lt;/a&gt;. Previously, the API buffered each tool call until its argument JSON completed — a design that added seconds of dead time per step in long agent chains.&lt;/p&gt;
&lt;h2&gt;What changed&lt;/h2&gt;
&lt;p&gt;With streaming enabled, agent frameworks receive partial argument deltas as the model generates them, in a format compatible with the OpenAI streaming convention most frameworks already parse. For K2 Thinking specifically — a model whose signature behavior is chains of hundreds of sequential tool calls — the latency saving compounds per step.&lt;/p&gt;
&lt;p&gt;The change is server-side. Model weights are unchanged; self-hosted deployments already had access to raw token streams and are unaffected.&lt;/p&gt;
&lt;h2&gt;Why it matters&lt;/h2&gt;
&lt;p&gt;Tool-call streaming was one of the last API-surface gaps between K2 Thinking and the closed-model APIs it competes with. Anthropic and OpenAI both ship fine-grained tool streaming as standard. With the gap closed, framework-level integrations no longer need Moonshot-specific buffering paths — one less reason for agent builders to default to closed APIs.&lt;/p&gt;
&lt;p&gt;K2 Thinking remains, in our tracked measurements, the open-weight model with the &lt;a href=&quot;/models/&quot;&gt;longest stable tool-call chains&lt;/a&gt;. API pricing is unchanged at $0.60 input / $2.50 output per million tokens (verified August 1, 2026).&lt;/p&gt;
</content:encoded><category>News</category><category>moonshot</category><category>kimi</category><category>api</category><category>agents</category></item><item><title>How we benchmark models on llm.blog</title><link>https://llm.blog/blog/how-we-benchmark/</link><guid isPermaLink="true">https://llm.blog/blog/how-we-benchmark/</guid><description>llm.blog measures models through a fixed public harness: vendor-default settings, three seeds per benchmark, median reported, every score dated. We re-measure monthly, flag unreproduced vendor claims as unverified, and publish configuration diffs whenever the harness changes. This page is the canonical methodology reference for every number on the leaderboard.</description><pubDate>Sat, 01 Aug 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;llm.blog measures models through a fixed public harness: vendor-default settings, three seeds per benchmark, median reported, every score dated. We re-measure monthly, flag unreproduced vendor claims as unverified, and publish configuration diffs whenever the harness changes. This page is the canonical methodology reference for every number on the leaderboard.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Every number on &lt;a href=&quot;/leaderboard/&quot;&gt;the leaderboard&lt;/a&gt; comes out of the process on this page. When this page and a leaderboard cell disagree, the cell&amp;#39;s measurement date tells you which is newer; the methodology described here is versioned, and harness changes are logged in the changelog with dates.&lt;/p&gt;
&lt;h2&gt;The five rules&lt;/h2&gt;
&lt;p&gt;Our measurement policy compresses to five rules:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Vendor-default settings.&lt;/strong&gt; We run each model the way a developer gets it from the API with no tuning: default temperature, default reasoning effort, vendor-recommended system prompt if the docs specify one. Benchmarks are meant to predict what you get, not what is achievable.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fixed public scaffold.&lt;/strong&gt; Every model runs the same agentless-lite scaffold for SWE-bench Verified and the official containers for Terminal-Bench. Scaffold code is pinned by commit hash and linked from each measurement cycle&amp;#39;s notes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Three seeds, median reported.&lt;/strong&gt; Single-run benchmark scores are noise wearing a suit. We run three seeds and report the median; the spread is retained internally and cited when a movement is within variance.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Every score carries a date.&lt;/strong&gt; Serving-stack revisions move scores without version-number changes. A score without a measurement date is not a fact; it is a rumor with decimals.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unreproduced numbers are flagged.&lt;/strong&gt; Where we display a vendor-published figure we have not reproduced — new model, new benchmark, harness gap — it renders with an explicit &lt;code&gt;unverified&lt;/code&gt; flag until we can measure it ourselves.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Benchmark selection&lt;/h2&gt;
&lt;p&gt;We track six benchmarks. The selection favors measures that predict production behavior for API practitioners over leaderboard-culture staples:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;What it predicts&lt;/th&gt;
&lt;th&gt;Known weakness&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;SWE-bench Verified&lt;/td&gt;
&lt;td&gt;Repository-scale coding agents&lt;/td&gt;
&lt;td&gt;Python-heavy; scaffold-sensitive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 2.0&lt;/td&gt;
&lt;td&gt;Operational/terminal competence&lt;/td&gt;
&lt;td&gt;Young benchmark, small task set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;Expert reasoning ceiling&lt;/td&gt;
&lt;td&gt;Saturating at the frontier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AIME 2025&lt;/td&gt;
&lt;td&gt;Structured mathematical reasoning&lt;/td&gt;
&lt;td&gt;Contamination pressure grows yearly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMMU&lt;/td&gt;
&lt;td&gt;Multimodal understanding&lt;/td&gt;
&lt;td&gt;Weak proxy for document extraction&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LiveCodeBench&lt;/td&gt;
&lt;td&gt;Contamination-controlled coding&lt;/td&gt;
&lt;td&gt;Contest style ≠ production code&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Two absences are deliberate. We do not track MMLU or its descendants: at frontier accuracy the residual questions are disproportionately mislabeled or ambiguous, so movement measures label noise. We do not track head-to-head preference Elo (arena-style rankings): preference voting measures persuasiveness under casual prompts, which diverges from task competence in ways that reward stylistic tuning.&lt;/p&gt;
&lt;h2&gt;Contamination policy&lt;/h2&gt;
&lt;p&gt;Contamination is not a scandal; it is a baseline condition to engineer around. Our posture:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Prefer benchmarks with rolling post-cutoff problem sets (LiveCodeBench) or verified human curation (SWE-bench Verified).&lt;/li&gt;
&lt;li&gt;Track year-stamped sets (AIME 2025) only while the newest models&amp;#39; training cutoffs plausibly precede the problems, and retire them after.&lt;/li&gt;
&lt;li&gt;Watch for the contamination signature — a model dramatically outperforming its own capability profile on one aged benchmark — and annotate rather than delete when we suspect it.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Variance and what counts as movement&lt;/h2&gt;
&lt;p&gt;Run-to-run spread on agentic benchmarks is larger than most coverage admits: ±0.4 points is typical on SWE-bench Verified at three seeds, and Terminal-Bench 2.0 runs wider. Our reporting rule: a month-over-month delta inside the observed seed spread is described as &amp;quot;stable,&amp;quot; not as movement. Deltas beyond spread get called movement and, where we can identify it, a cause — serving update, harness revision, or unexplained.&lt;/p&gt;
&lt;p&gt;When the harness itself changes — container image updates, scaffold pin bumps — we re-run affected benchmarks for &lt;strong&gt;all&lt;/strong&gt; models in the same cycle, never just the new arrivals, and note the change in the cycle notes. A harness change that moves any median more than 0.5 points triggers a full historical annotation.&lt;/p&gt;
&lt;h2&gt;What the leaderboard is not&lt;/h2&gt;
&lt;p&gt;The leaderboard ranks measured capability on six axes. It does not rank:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Value.&lt;/strong&gt; Price-per-capability is workload-specific arithmetic; the &lt;a href=&quot;/compare/&quot;&gt;comparison pages&lt;/a&gt; do that math per pairing, and the calculator on each one does it for your token mix.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Safety or alignment properties.&lt;/strong&gt; We do not have the harness to measure them credibly, and we decline to launder vibes into a column.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Latency and throughput.&lt;/strong&gt; Serving performance varies by region, tier, and hour; point-in-time measurements would mislead more than inform. We may add percentile-tracked latency in future with continuous measurement.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Reproducing our numbers&lt;/h2&gt;
&lt;p&gt;Scaffold commit hashes, container digests, seed values, and per-task result JSON for each cycle are retained. The dataset behind the leaderboard — scores, dates, sources, flags — is the same JSON that renders the site, and the &lt;a href=&quot;/llms.txt&quot;&gt;llms.txt&lt;/a&gt; file points machine readers at the canonical URLs. Cite the leaderboard with its measurement date and the numbers will stay attributable even after later refreshes move them.&lt;/p&gt;
&lt;p&gt;Corrections: &lt;a href=&quot;mailto:corrections@llm.blog&quot;&gt;corrections@llm.blog&lt;/a&gt;. A correction that survives verification lands in the next cycle notes with credit.&lt;/p&gt;
</content:encoded><category>Deep Dives</category><category>methodology</category><category>benchmarks</category><category>leaderboard</category></item><item><title>OpenAI halves GPT-5.1 cached-input pricing to $0.125 per million tokens</title><link>https://llm.blog/news/gpt-5-1-cached-input-price-cut/</link><guid isPermaLink="true">https://llm.blog/news/gpt-5-1-cached-input-price-cut/</guid><description>OpenAI reduced GPT-5.1 cached-input pricing from $0.25 to $0.125 per million tokens on July 28, 2026 — 10% of the $1.25 base input rate. Base input and output prices are unchanged. Workloads reusing long system prompts, tool definitions, or document contexts see the largest savings.</description><pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;OpenAI reduced GPT-5.1 cached-input pricing from $0.25 to $0.125 per million tokens on July 28, 2026 — 10% of the $1.25 base input rate. Base input and output prices are unchanged. Workloads reusing long system prompts, tool definitions, or document contexts see the largest savings.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;OpenAI cut GPT-5.1 cached-input pricing to $0.125 per million tokens on July 28, 2026, down from $0.25, per the &lt;a href=&quot;https://platform.openai.com/docs/pricing&quot;&gt;pricing documentation&lt;/a&gt;. Cached input now bills at 10% of the $1.25 base input rate. Base input and output ($10 per million) prices are unchanged.&lt;/p&gt;
&lt;h2&gt;Who saves&lt;/h2&gt;
&lt;p&gt;Prompt caching applies to repeated prefixes: system prompts, tool schemas, few-shot blocks, and pinned document context. The workloads that benefit most:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agents with large tool manifests.&lt;/strong&gt; A 20K-token tool schema reused across a 50-step session now bills the repeated portion at $0.125 per million — an 87.5% discount against uncached input.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;High-frequency RAG over stable corpora.&lt;/strong&gt; Pipelines that pin document context across many queries convert most of their input volume to the cached rate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multi-tenant chat products&lt;/strong&gt; with shared system prompts across users on the same cache shard.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Cache TTL and eligibility rules are unchanged; only the rate moved.&lt;/p&gt;
&lt;h2&gt;Competitive context&lt;/h2&gt;
&lt;p&gt;The cut narrows the effective-price gap with Claude Sonnet 4.5, whose cache-read pricing has been $0.30 per million tokens (10% of its $3 base) since launch. Both vendors now anchor cached input at 10% of base — the emerging industry convention — leaving the base rates as the real comparison. Current effective prices for common token mixes are in our &lt;a href=&quot;/compare/&quot;&gt;GPT-5.1 vs Claude Sonnet 4.5 comparison&lt;/a&gt;, updated July 28, 2026.&lt;/p&gt;
</content:encoded><category>News</category><category>pricing</category><category>openai</category><category>gpt</category></item><item><title>Alibaba ships Qwen3-Max 2026-07 update with improved long-context recall</title><link>https://llm.blog/news/qwen3-max-july-update/</link><guid isPermaLink="true">https://llm.blog/news/qwen3-max-july-update/</guid><description>Alibaba released a Qwen3-Max update on July 22, 2026, improving retrieval accuracy in the second half of its 256K-token context window. The update is transparent to API callers — same model ID, same tiered pricing from $1.20 per million input tokens. Alibaba reports no changes to reasoning benchmark scores.</description><pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Alibaba released a Qwen3-Max update on July 22, 2026, improving retrieval accuracy in the second half of its 256K-token context window. The update is transparent to API callers — same model ID, same tiered pricing from $1.20 per million input tokens. Alibaba reports no changes to reasoning benchmark scores.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Alibaba deployed an updated Qwen3-Max on July 22, 2026, per the &lt;a href=&quot;https://qwen.ai/blog&quot;&gt;Qwen team blog&lt;/a&gt;. The update targets long-context retrieval: Alibaba reports materially improved needle-in-haystack and multi-hop recall between 128K and 256K tokens, the region where the model previously degraded fastest.&lt;/p&gt;
&lt;h2&gt;What changed&lt;/h2&gt;
&lt;p&gt;The update is a serving-side model revision under the same &lt;code&gt;qwen3-max&lt;/code&gt; API identifier — callers get it automatically. Alibaba states reasoning and coding benchmark scores are unchanged, and our tracked scores for the model keep their existing measurement dates until the August refresh confirms.&lt;/p&gt;
&lt;p&gt;Tiered pricing is unchanged, starting at $1.20 per million input tokens and $6 per million output for the first tier (verified July 22, 2026, &lt;a href=&quot;https://www.alibabacloud.com/help/en/model-studio/models&quot;&gt;Model Studio documentation&lt;/a&gt;).&lt;/p&gt;
&lt;h2&gt;Practitioner notes&lt;/h2&gt;
&lt;p&gt;Long-context claims are the least transferable numbers in vendor announcements: recall curves depend heavily on task shape, and vendor-reported needle tests overstate practical retrieval in structured-extraction workloads. We flag Qwen3-Max&amp;#39;s long-context figures &lt;code&gt;unverified&lt;/code&gt; on &lt;a href=&quot;/models/&quot;&gt;its model page&lt;/a&gt; until the September measurement cycle, which will add a long-context retrieval track to &lt;a href=&quot;/blog/how-we-benchmark/&quot;&gt;our methodology&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Teams already running Qwen3-Max on document-heavy workloads should A/B the update against archived outputs — same-ID serving revisions are exactly the case where silent behavior drift bites production pipelines.&lt;/p&gt;
</content:encoded><category>News</category><category>alibaba</category><category>qwen</category><category>long-context</category></item><item><title>Structured-output verification patterns for production LLM pipelines</title><link>https://llm.blog/blog/structured-output-verification/</link><guid isPermaLink="true">https://llm.blog/blog/structured-output-verification/</guid><description>Schema-constrained decoding eliminates malformed JSON but not wrong JSON: type-valid hallucinated values, unit drift, and cross-field contradictions survive it. Production pipelines need layered verification — schema, semantic invariants, provenance grounding, statistical monitoring, and selective human escalation. This piece documents each layer with implementation patterns and observed failure rates.</description><pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Schema-constrained decoding eliminates malformed JSON but not wrong JSON: type-valid hallucinated values, unit drift, and cross-field contradictions survive it. Production pipelines need layered verification — schema, semantic invariants, provenance grounding, statistical monitoring, and selective human escalation. This piece documents each layer with implementation patterns and observed failure rates.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Every provider now ships schema-constrained decoding, and every provider&amp;#39;s documentation quietly overpromises what it buys you. Constrained decoding guarantees the output &lt;em&gt;parses&lt;/em&gt;. It does not guarantee the output is &lt;em&gt;true&lt;/em&gt;, &lt;em&gt;grounded&lt;/em&gt;, or &lt;em&gt;internally consistent&lt;/em&gt;. The gap between those guarantees is where production incidents live.&lt;/p&gt;
&lt;p&gt;This piece documents the verification stack we run on a document-extraction pipeline processing roughly 40K documents daily against frontier-model APIs, with the failure rates each layer catches. Numbers are from June–July 2026 production traffic; your distribution will differ — measure it.&lt;/p&gt;
&lt;h2&gt;The failure taxonomy schema validation can&amp;#39;t see&lt;/h2&gt;
&lt;p&gt;Schema-valid failures cluster into four families:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Type-valid hallucination.&lt;/strong&gt; The model returns &lt;code&gt;&amp;quot;invoice_total&amp;quot;: 4820.00&lt;/code&gt; — a perfectly typed float that appears nowhere in the document. Constrained decoding &lt;em&gt;increases&lt;/em&gt; the risk at the margin: when the model is uncertain, the grammar forbids it from expressing uncertainty outside the schema.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unit and format drift.&lt;/strong&gt; Dates parsed day-first from US documents, totals in cents where the schema assumes dollars, percentages as &lt;code&gt;0.15&lt;/code&gt; vs &lt;code&gt;15&lt;/code&gt;. Type-correct, silently wrong by 100×.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-field contradiction.&lt;/strong&gt; Line items that don&amp;#39;t sum to the stated total; an end date before a start date; a &lt;code&gt;currency: EUR&lt;/code&gt; beside a &lt;code&gt;$&lt;/code&gt;-prefixed source span.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Enum coercion.&lt;/strong&gt; Forced to choose from &lt;code&gt;[&amp;quot;invoice&amp;quot;, &amp;quot;receipt&amp;quot;, &amp;quot;credit_note&amp;quot;]&lt;/code&gt;, the model files a purchase order as an invoice rather than failing. Closed enums without an escape value convert &amp;quot;I don&amp;#39;t know&amp;quot; into confident misclassification.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In our pipeline, 2–4% of schema-valid extractions exhibit at least one of these per day, varying with document mix.&lt;/p&gt;
&lt;h2&gt;Layer 1: schema — but design for refusal&lt;/h2&gt;
&lt;p&gt;The schema layer is table stakes; the design detail that matters is giving the model somewhere to put uncertainty:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &amp;quot;doc_type&amp;quot;: { &amp;quot;enum&amp;quot;: [&amp;quot;invoice&amp;quot;, &amp;quot;receipt&amp;quot;, &amp;quot;credit_note&amp;quot;, &amp;quot;other&amp;quot;] },
  &amp;quot;confidence&amp;quot;: { &amp;quot;enum&amp;quot;: [&amp;quot;high&amp;quot;, &amp;quot;medium&amp;quot;, &amp;quot;low&amp;quot;] },
  &amp;quot;fields_not_found&amp;quot;: { &amp;quot;type&amp;quot;: &amp;quot;array&amp;quot;, &amp;quot;items&amp;quot;: { &amp;quot;type&amp;quot;: &amp;quot;string&amp;quot; } }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Adding &lt;code&gt;&amp;quot;other&amp;quot;&lt;/code&gt; plus a &lt;code&gt;fields_not_found&lt;/code&gt; array cut our enum-coercion rate by more than half. The model was never unable to say &amp;quot;I don&amp;#39;t know&amp;quot; — the schema had just made it grammatically impossible.&lt;/p&gt;
&lt;p&gt;Nullable-with-reason beats required-with-guess for every field a document might legitimately lack.&lt;/p&gt;
&lt;h2&gt;Layer 2: semantic invariants&lt;/h2&gt;
&lt;p&gt;Every schema ships with unstated physics. Write them down as executable checks that run on every extraction:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-python&quot;&gt;def invariants(x: Extraction) -&amp;gt; list[Violation]:
    v = []
    if x.line_items and abs(sum(i.amount for i in x.line_items) - x.total) &amp;gt; 0.01:
        v.append(Violation(&amp;quot;line_items_sum&amp;quot;, severity=&amp;quot;block&amp;quot;))
    if x.due_date and x.issue_date and x.due_date &amp;lt; x.issue_date:
        v.append(Violation(&amp;quot;date_order&amp;quot;, severity=&amp;quot;block&amp;quot;))
    if x.total and not (0 &amp;lt; x.total &amp;lt; 10_000_000):
        v.append(Violation(&amp;quot;total_range&amp;quot;, severity=&amp;quot;review&amp;quot;))
    return v
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Invariants are cheap, deterministic, and catch the contradiction family almost completely. Ours run in three severities: &lt;code&gt;block&lt;/code&gt; (re-extract with violation appended to the prompt), &lt;code&gt;review&lt;/code&gt; (human queue), &lt;code&gt;log&lt;/code&gt; (monitoring only). The re-extract-with-feedback loop resolves roughly 70% of &lt;code&gt;block&lt;/code&gt; violations on the first retry; the remainder queue for review.&lt;/p&gt;
&lt;p&gt;The prompt-side counterpart matters as much: state the invariants in the extraction prompt. Models violate constraints they were never told about at several times the rate of stated ones.&lt;/p&gt;
&lt;h2&gt;Layer 3: provenance grounding&lt;/h2&gt;
&lt;p&gt;Type-valid hallucination survives layers 1 and 2 when the invented value is plausible. The countermeasure is requiring the model to cite its work — every extracted value paired with a source span:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{ &amp;quot;total&amp;quot;: 4820.00, &amp;quot;total_span&amp;quot;: &amp;quot;TOTAL DUE: $4,820.00&amp;quot; }
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Verification then checks the span actually occurs in the source (fuzzy match tolerating OCR noise and whitespace) and the value is derivable from it (number extraction on the span, comparison against the field). Span-not-found or value-mismatch routes to review.&lt;/p&gt;
&lt;p&gt;Grounding catches the failures that matter most — invented financial values — at a cost: span fields roughly double output tokens for dense schemas. We ground only the six fields whose wrongness is expensive, not all forty. Failure rate at this layer: ~0.5% of extractions that passed layers 1–2, which is exactly the population you could not have found otherwise.&lt;/p&gt;
&lt;h2&gt;Layer 4: statistical monitoring&lt;/h2&gt;
&lt;p&gt;Per-document verification misses distribution-level drift: a serving-stack update that shifts date formats, a new document template that quietly halves field-found rates. Distribution monitors watch daily aggregates per field — null rate, mean, P95, enum mix — against trailing 28-day bands, alerting on excursions.&lt;/p&gt;
&lt;p&gt;This is the layer that catches &lt;em&gt;silent model updates&lt;/em&gt;. Same API identifier, new serving revision, moved behavior: we have caught three such shifts in a year of operation, each visible in field-level distributions days before any per-document check fired. If your pipeline pins vendor aliases rather than dated snapshots, this layer is not optional.&lt;/p&gt;
&lt;h2&gt;Layer 5: selective human escalation&lt;/h2&gt;
&lt;p&gt;Escalation is a budget allocation problem: review capacity is fixed, so route it where per-item expected loss is highest. Our routing score is a weighted sum of layer signals — invariant &lt;code&gt;review&lt;/code&gt; flags, grounding mismatches, model-reported low confidence, and monetary value of the document. The result: ~1.5% of daily volume queues for human review, and that 1.5% contains an estimated 80%+ of surviving errors (estimated via periodic random-sample audits — which you must also run, or your escalation model grades its own homework).&lt;/p&gt;
&lt;h2&gt;The stack, summarized&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Catches&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Our observed catch rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Schema + refusal design&lt;/td&gt;
&lt;td&gt;Malformed output, forced guessing&lt;/td&gt;
&lt;td&gt;~0&lt;/td&gt;
&lt;td&gt;eliminates class&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic invariants&lt;/td&gt;
&lt;td&gt;Contradictions, range violations&lt;/td&gt;
&lt;td&gt;CPU-trivial&lt;/td&gt;
&lt;td&gt;2–4% of valid outputs flagged&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provenance grounding&lt;/td&gt;
&lt;td&gt;Plausible hallucination&lt;/td&gt;
&lt;td&gt;~2× output tokens on grounded fields&lt;/td&gt;
&lt;td&gt;~0.5% post-invariant&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distribution monitoring&lt;/td&gt;
&lt;td&gt;Silent drift, template shifts&lt;/td&gt;
&lt;td&gt;infra-light&lt;/td&gt;
&lt;td&gt;3 serving shifts / year&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Selective escalation&lt;/td&gt;
&lt;td&gt;Residual tail&lt;/td&gt;
&lt;td&gt;fixed human budget&lt;/td&gt;
&lt;td&gt;~80% of survivors in 1.5% of volume&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;None of these layers is novel; the system property comes from running all five. Schema validation alone is a seatbelt bolted to a car with no brakes — it makes the crash tidier, not less likely.&lt;/p&gt;
&lt;h2&gt;What we&amp;#39;d build differently today&lt;/h2&gt;
&lt;p&gt;Two revisions from a year of operation. First, we would adopt grounded spans from day one rather than retrofitting them after the first hallucinated-total incident; retrofit cost exceeded first-build cost by an embarrassing multiple. Second, we would version extraction prompts and schemas in the same artifact with the same review gates as code — the worst drift we shipped came from a &amp;quot;harmless&amp;quot; prompt wording change that moved enum distributions 9%, discovered by layer 4 eleven days later. Everything upstream of the model call is code; treat it with code&amp;#39;s discipline.&lt;/p&gt;
</content:encoded><category>Deep Dives</category><category>engineering</category><category>structured-output</category><category>pipelines</category><category>reliability</category></item><item><title>Terminal-Bench 2.0 joins the llm.blog tracked benchmark set</title><link>https://llm.blog/news/terminal-bench-2-added/</link><guid isPermaLink="true">https://llm.blog/news/terminal-bench-2-added/</guid><description>llm.blog added Terminal-Bench 2.0 to its tracked benchmark set on July 15, 2026, scoring six models in the first cycle. Claude Opus 4.5 leads at 59.3%. Terminal-Bench measures task completion in a live terminal environment — operational competence that SWE-bench&apos;s repository-patching format does not capture.</description><pubDate>Wed, 15 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;llm.blog added Terminal-Bench 2.0 to its tracked benchmark set on July 15, 2026, scoring six models in the first cycle. Claude Opus 4.5 leads at 59.3%. Terminal-Bench measures task completion in a live terminal environment — operational competence that SWE-bench&apos;s repository-patching format does not capture.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;We added Terminal-Bench 2.0 to the tracked benchmark set on July 15, 2026, and scored six models in the first measurement cycle: Claude Opus 4.5 (59.3%), Gemini 3 Pro (54.2%), Claude Sonnet 4.5 (50.0%), Kimi K2 Thinking (47.1%), Claude Haiku 4.5 (41.0%), and Grok 4.1 (pending a harness fix). Remaining models join in August. All scores and dates are on the &lt;a href=&quot;/leaderboard/&quot;&gt;leaderboard&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Why Terminal-Bench&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://www.tbench.ai/&quot;&gt;Terminal-Bench&lt;/a&gt; drops a model into a live terminal and scores task completion: environment setup, debugging, data wrangling, service configuration. It measures operational competence — the &amp;quot;can it actually drive a computer&amp;quot; question — where SWE-bench Verified measures repository-scale code modification. The two disagree in informative ways: models tuned aggressively for patch generation can underperform on open-ended terminal tasks, and the gap predicts real-world agent reliability better than either number alone.&lt;/p&gt;
&lt;h2&gt;Configuration&lt;/h2&gt;
&lt;p&gt;We run the standard 2.0 task set in the official container, vendor-default reasoning settings, 20-minute per-task timeout, two attempts with the better run discarded (we report first-attempt scores; the second run measures variance only). Full harness details are in &lt;a href=&quot;/blog/how-we-benchmark/&quot;&gt;the methodology&lt;/a&gt;, which has been updated with the Terminal-Bench section as of July 15, 2026.&lt;/p&gt;
&lt;p&gt;Vendor-published Terminal-Bench numbers exist for several tracked models and differ from ours by up to four points, mostly attributable to scaffold differences. Model pages flag vendor figures &lt;code&gt;unverified&lt;/code&gt; where we have not reproduced them.&lt;/p&gt;
</content:encoded><category>News</category><category>benchmarks</category><category>methodology</category><category>terminal-bench</category></item><item><title>DeepSeek extends off-peak API discount window to 12 hours daily</title><link>https://llm.blog/news/deepseek-off-peak-window/</link><guid isPermaLink="true">https://llm.blog/news/deepseek-off-peak-window/</guid><description>DeepSeek extended its off-peak API discount window to 12 hours daily on July 8, 2026, up from 8. Requests between 16:30 and 04:30 UTC bill at roughly 50% of standard rates — about $0.14 input and $0.21 output per million tokens for DeepSeek-V3.2. Batch workloads gain the most.</description><pubDate>Wed, 08 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;DeepSeek extended its off-peak API discount window to 12 hours daily on July 8, 2026, up from 8. Requests between 16:30 and 04:30 UTC bill at roughly 50% of standard rates — about $0.14 input and $0.21 output per million tokens for DeepSeek-V3.2. Batch workloads gain the most.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;DeepSeek widened its off-peak discount window to 12 hours daily on July 8, 2026, per the &lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;API pricing page&lt;/a&gt;. Requests submitted between 16:30 and 04:30 UTC now bill at approximately half of standard rates — for DeepSeek-V3.2, roughly $0.14 per million input tokens and $0.21 per million output, against the standard $0.28 / $0.42.&lt;/p&gt;
&lt;h2&gt;The arbitrage&lt;/h2&gt;
&lt;p&gt;Off-peak pricing rewards workloads that can shift: nightly embedding refreshes, evaluation sweeps, synthetic-data generation, batch summarization. A pipeline moving 5B tokens of monthly batch volume into the window saves about $700 per month at V3.2 rates — on prices that were already the lowest in our tracked set.&lt;/p&gt;
&lt;p&gt;The mechanics are simple: pricing is determined by request submission time, no reservation or commitment required. That makes the window one config change for most batch schedulers.&lt;/p&gt;
&lt;h2&gt;The bigger pattern&lt;/h2&gt;
&lt;p&gt;Time-of-day pricing is DeepSeek exporting its infrastructure reality: its GPU fleet&amp;#39;s demand curve follows Chinese business hours, and the discount window sells the trough. No Western lab has followed — OpenAI, Anthropic, and Google price batch workloads through dedicated batch APIs with 50% discounts and delayed completion instead. The two mechanisms converge on similar effective rates; DeepSeek&amp;#39;s version returns results synchronously, which matters for evaluation loops that gate on results.&lt;/p&gt;
&lt;p&gt;Standard-rate comparisons across all tracked models are on the &lt;a href=&quot;/leaderboard/&quot;&gt;leaderboard&lt;/a&gt;; the &lt;a href=&quot;/models/&quot;&gt;DeepSeek-V3.2 page&lt;/a&gt; carries the full pricing history.&lt;/p&gt;
</content:encoded><category>News</category><category>pricing</category><category>deepseek</category><category>api</category></item><item><title>Claude Sonnet 4.5&apos;s 1M-token context window graduates from beta</title><link>https://llm.blog/news/sonnet-1m-context-ga/</link><guid isPermaLink="true">https://llm.blog/news/sonnet-1m-context-ga/</guid><description>Anthropic moved Claude Sonnet 4.5&apos;s 1M-token context window from beta to general availability on July 2, 2026. The beta header requirement is removed. Prompts above 200K tokens bill at premium rates — $6 input and $22.50 output per million tokens — doubling effective input cost for full-window workloads.</description><pubDate>Thu, 02 Jul 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;Anthropic moved Claude Sonnet 4.5&apos;s 1M-token context window from beta to general availability on July 2, 2026. The beta header requirement is removed. Prompts above 200K tokens bill at premium rates — $6 input and $22.50 output per million tokens — doubling effective input cost for full-window workloads.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;Anthropic promoted Claude Sonnet 4.5&amp;#39;s 1M-token context window to general availability on July 2, 2026, per the &lt;a href=&quot;https://docs.anthropic.com/en/docs/build-with-claude/context-windows&quot;&gt;context windows documentation&lt;/a&gt;. The &lt;code&gt;context-1m&lt;/code&gt; beta header is no longer required; all API traffic can address the full window.&lt;/p&gt;
&lt;h2&gt;The pricing structure&lt;/h2&gt;
&lt;p&gt;The window is tiered. Prompts up to 200K tokens bill at the standard $3 input / $15 output per million. Prompts above 200K bill the entire request at the long-context rate of $6 input / $22.50 output per million tokens (rates verified July 2, 2026).&lt;/p&gt;
&lt;p&gt;That cliff shape matters for pipeline design: a 210K-token prompt costs more than twice a 195K-token prompt on input. Retrieval pipelines that can compress below the boundary should; the premium tier is priced for workloads that genuinely need single-pass reasoning over massive context — codebase-scale analysis, long-document synthesis, extended agent sessions with full history.&lt;/p&gt;
&lt;h2&gt;Competitive position&lt;/h2&gt;
&lt;p&gt;GA closes most of the context gap with Gemini 3 Pro, whose 1M-token window has been standard since &lt;a href=&quot;https://blog.google/products/gemini/gemini-3/&quot;&gt;its November 2025 launch&lt;/a&gt; with a gentler tier boundary ($4 input above 200K vs Anthropic&amp;#39;s $6). For pure long-document work, Gemini remains cheaper at the top of the window; Sonnet&amp;#39;s case rests on its agentic-coding lead. Both patterns are priced out in &lt;a href=&quot;/compare/&quot;&gt;our head-to-head&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;/models/&quot;&gt;Claude Sonnet 4.5 model page&lt;/a&gt; reflects GA status in its spec table as of July 2, 2026.&lt;/p&gt;
</content:encoded><category>News</category><category>anthropic</category><category>claude</category><category>long-context</category><category>pricing</category></item></channel></rss>