Prompt caching is a feature offered by Anthropic, OpenAI, and other AI providers that lets a model reuse the processing it already did on a repeated block of text instead of recomputing it from scratch on every request. In practice, that means you can send the same long system prompt, document, or set of tool definitions on request after request and pay a fraction of the normal price for the tokens that haven’t changed — while also getting a faster response.
How Does Prompt Caching Actually Work?
Every time you call a large language model, the model has to convert your prompt into an internal representation before it can generate a reply. When most of that prompt is identical to a previous call — the same system instructions, the same reference document, the same few-shot examples — recomputing it every single time is wasted work. Prompt caching stores that internal representation after the first request so a later request that reuses the same prefix can skip straight to processing only the new part.
With Claude models, you mark where the reusable part of your prompt ends using a cache_control parameter, either letting the system apply a single automatic breakpoint or setting up to four manual breakpoints for finer control. On the next request, Claude checks whether that prefix is already cached and, if so, reuses it instead of reprocessing it.
OpenAI’s approach is automatic: for prompts above a minimum length, the API detects repeated prefixes on its own and applies the discount without any extra parameters, reusing the saved key-value state from the earlier request.
How Much Money Does Prompt Caching Actually Save?
The numbers are substantial enough to change how teams design their prompts. On Claude, reading from the cache costs roughly 0.1x the base input price — about a 90% discount — while writing to the 5-minute cache costs 1.25x the base price and writing to the optional 1-hour cache costs 2x. So the first call that populates the cache costs slightly more, but every call afterward that reuses it costs dramatically less.
OpenAI advertises a similar figure: cached input tokens are billed at a discount of up to 90% compared to standard input pricing, applied automatically once a prompt crosses the minimum cacheable length.
- Claude cache reads: about 0.1x base input price (≈90% cheaper)
- Claude 5-minute cache writes: 1.25x base input price
- Claude 1-hour cache writes: 2x base input price
- OpenAI cached input tokens: up to 90% cheaper than standard input pricing
What’s the Minimum Prompt Length and Cache Lifetime?
Caching only kicks in once a prompt is long enough to be worth storing. For most current Claude models, including Claude Sonnet, the minimum is 1,024 tokens; some larger models drop that threshold to 512 tokens. Below that, the API simply processes the request normally without an error.
Cache lifetime matters just as much as the discount. Claude’s default cache lasts 5 minutes from the start of the request, with an optional 1-hour cache for prompts reused less frequently than every few minutes but more often than hourly. OpenAI’s cached entries typically stay available for 5 to 30 minutes of inactivity depending on the model, with extended retention options for higher-volume use cases.
Prompt Caching vs. Prompt Compression: What’s the Difference?
It’s easy to confuse prompt caching with prompt compression, but they solve different problems. Compression shrinks the prompt itself — removing redundant wording so every request, cached or not, uses fewer tokens. Caching doesn’t change the size of your prompt at all; it changes whether the model has to reprocess that size from zero on every call. The two techniques stack well together: a compressed prompt that’s also cached is both smaller and cheaper to reuse.
When Should You Actually Use Prompt Caching?
Prompt caching pays off most clearly in a handful of common patterns: long system prompts that stay fixed across many user turns, multi-turn conversations where earlier messages are resent with every new call, RAG pipelines that inject the same reference documents repeatedly, coding assistants that keep a large codebase or style guide in context, and any workflow with a large, static set of tool definitions or few-shot examples. If your prompt changes substantially on every call, caching won’t help much — the value comes from repetition, not from any single request.
Frequently Asked Questions
Is prompt caching the same as fine-tuning a model?
No. Fine-tuning permanently changes a model’s weights based on training examples. Prompt caching is temporary and per-request: it stores the processed version of a specific prompt prefix for a few minutes to an hour so future calls with that same prefix skip redundant computation. The model itself never changes.
Do I lose any output quality by using prompt caching?
No. Caching only affects how the prompt is processed internally before generation starts — it doesn’t change the model’s weights, the content of your prompt, or the reasoning it applies to produce a response. The output should be equivalent to an uncached call with the same prompt.
How do I know if caching is actually being used on my requests?
Both providers return usage fields in the API response that report cache activity. Claude’s response includes cache_creation_input_tokens and cache_read_input_tokens, so you can confirm a cache write or a cache hit occurred on any given call and track your savings directly from the API response.
Sources: Anthropic – Prompt caching documentation, OpenAI – Prompt caching guide.



