OpenAI halves GPT-5.1 cached-input pricing to $0.125 per million tokens
Draft — pending editorial review
OpenAI reduced GPT-5.1 cached-input pricing from $0.25 to $0.125 per million tokens on July 28, 2026 — 10% of the $1.25 base input rate. Base input and output prices are unchanged. Workloads reusing long system prompts, tool definitions, or document contexts see the largest savings.
OpenAI cut GPT-5.1 cached-input pricing to $0.125 per million tokens on July 28, 2026, down from $0.25, per the pricing documentation. Cached input now bills at 10% of the $1.25 base input rate. Base input and output ($10 per million) prices are unchanged.
Who saves
Prompt caching applies to repeated prefixes: system prompts, tool schemas, few-shot blocks, and pinned document context. The workloads that benefit most:
- Agents with large tool manifests. A 20K-token tool schema reused across a 50-step session now bills the repeated portion at $0.125 per million — an 87.5% discount against uncached input.
- High-frequency RAG over stable corpora. Pipelines that pin document context across many queries convert most of their input volume to the cached rate.
- Multi-tenant chat products with shared system prompts across users on the same cache shard.
Cache TTL and eligibility rules are unchanged; only the rate moved.
Competitive context
The cut narrows the effective-price gap with Claude Sonnet 4.5, whose cache-read pricing has been $0.30 per million tokens (10% of its $3 base) since launch. Both vendors now anchor cached input at 10% of base — the emerging industry convention — leaving the base rates as the real comparison. Current effective prices for common token mixes are in our GPT-5.1 vs Claude Sonnet 4.5 comparison, updated July 28, 2026.