Skip to content

News · pricing · openai · gpt

OpenAI halves GPT-5.1 cached-input pricing to $0.125 per million tokens

Draft — pending editorial review

OpenAI reduced GPT-5.1 cached-input pricing from $0.25 to $0.125 per million tokens on July 28, 2026 — 10% of the $1.25 base input rate. Base input and output prices are unchanged. Workloads reusing long system prompts, tool definitions, or document contexts see the largest savings.

By Kushagra Sikka1 min read

OpenAI cut GPT-5.1 cached-input pricing to $0.125 per million tokens on July 28, 2026, down from $0.25, per the pricing documentation. Cached input now bills at 10% of the $1.25 base input rate. Base input and output ($10 per million) prices are unchanged.

Who saves

Prompt caching applies to repeated prefixes: system prompts, tool schemas, few-shot blocks, and pinned document context. The workloads that benefit most:

  • Agents with large tool manifests. A 20K-token tool schema reused across a 50-step session now bills the repeated portion at $0.125 per million — an 87.5% discount against uncached input.
  • High-frequency RAG over stable corpora. Pipelines that pin document context across many queries convert most of their input volume to the cached rate.
  • Multi-tenant chat products with shared system prompts across users on the same cache shard.

Cache TTL and eligibility rules are unchanged; only the rate moved.

Competitive context

The cut narrows the effective-price gap with Claude Sonnet 4.5, whose cache-read pricing has been $0.30 per million tokens (10% of its $3 base) since launch. Both vendors now anchor cached input at 10% of base — the emerging industry convention — leaving the base rates as the real comparison. Current effective prices for common token mixes are in our GPT-5.1 vs Claude Sonnet 4.5 comparison, updated July 28, 2026.