Large Language ModelsGenerate imagesGenerate videos
Prompt Caching Explained: Claude, OpenAI and Gemini Compared
Prompt caching can cut the cost of repeated input to a fraction, but Claude, OpenAI and Gemini each run it differently. See how breakpoints, lifetimes, write fees and minimum sizes compare, with a worked cost example and the mistakes that stop cache hits.
Every request you send to a large language model starts with the same expensive ritual. The system prompt, the tool definitions, the pasted documentation and the few-shot examples all get read again, token by token, at full price, even though nothing in them has changed since the last call. Prompt caching ends that waste. The provider stores the processed form of your repeated prefix and charges a fraction of the normal input price when the next request reuses it. Think of a restaurant's mise en place: the cook chops the shallots once before service, and every order after that is assembled from ready containers.
Claude, OpenAI and Gemini all offer the idea, but they disagree on almost everything else: who decides what gets cached, how long it lives, what it costs to write, and how big a prompt must be before it qualifies. This article lines the three up using numbers from their current documentation, so you can pick a setup and dodge the mistakes that quietly switch caching off.
What Prompt Caching Actually Does
When a model reads your prompt, it does real work for every token before it can write the first word of the answer. That work is identical every time the beginning of the prompt is identical. Prompt caching lets the provider keep the result of that work for a short window and reuse it, so the repeated part is processed once and billed at a discount afterward.
Picture a card catalog. The provider files the processed prefix under a fingerprint of its exact content. The next request that produces the same fingerprint pulls the drawer instead of rebuilding the contents from scratch.
The Prefix Rule
Caching works from the first token forward. The cached part has to be an exact, unbroken match for the start of your request, and the first difference ends the match. Everything after that point is processed and billed at the regular rate.
That single rule shapes every provider's advice. Put the content that never changes at the top, and push everything that changes (the user's question, the current date, retrieved snippets) to the bottom. A good top section usually holds:
Tool definitions that stay the same between calls
System instructions and style rules
Reference material such as a manual, a contract or a codebase summary
Few-shot examples you reuse on every request
💡 Quick check: if you printed two requests and highlighted what they share, the highlighted part must be one solid block that begins at the very first character. A shared paragraph in the middle does not count.
Where the Savings Come From
Three things improve when the prefix is reused:
Cost: cached tokens bill at a small fraction of the normal input rate, while only the new tail is charged in full.
Latency: a long prompt means a longer wait before the first output token. Skipping the reprocessing shortens that wait, and the gain grows with the length of the prefix.
Rate limits: Anthropic states that cache reads are not deducted from your rate limit, so cached traffic leaves room for more requests.
Agent loops, document question answering, long chat histories and classifiers with a big rubric benefit the most. A prompt that is different every time gets nothing.
Claude: Explicit Control
Claude gives you the most direct control of the three. You decide where the cacheable prefix ends, and you decide how long it should live.
Breakpoints and Automatic Mode
You mark a content block with cache_control of type ephemeral, and everything from the start of the request up to and including that block becomes the cached prefix. The order is fixed: tools first, then system, then messages. You can set up to four breakpoints, and the API returns a 400 error if you try to place a fifth block-level one.
There is also an automatic mode. Add a single cache_control field at the top level of the request, and the system applies the breakpoint to the last cacheable block, then moves it forward as the conversation grows.
One detail catches long agent sessions. When the system looks for an earlier matching entry, it checks at most 20 positions back from each breakpoint. If a turn adds many blocks, the previous entry can fall outside that window, and the fix is a second breakpoint placed earlier in the prompt.
Prices and Lifetimes
The default lifetime is 5 minutes, and every hit refreshes the timer at no charge. If your traffic comes in slow bursts, you can opt into one hour with "ttl": "1h" at a higher write price.
Item
Multiplier
Opus 5.5 example (per million tokens)
Base input
1x
$4.00
5-minute cache write
1.25x
$5.00
1-hour cache write
2x
$8.00
Cache read
0.05x on this model
$0.20
Most Claude models read from the cache at 0.1x of the base input price. The Opus 5.5 and Sonnet 5.5 tiers read at 0.05x, and some newer tiers go lower still.
Minimum Sizes by Model
The threshold depends on the model, and it has moved a lot between generations:
512 tokens: the 5.x generation, including Claude Fable 5
💡 Silent failure: a prompt shorter than the minimum is processed without caching, and no error is returned. The only sign is a cache field that stays at zero.
OpenAI: Breakpoints Arrive
OpenAI's story has two chapters. For a long time caching was fully automatic. With GPT-5.6 it gained breakpoints and a write fee, which makes it look much more like Claude.
Older Models: Fully Automatic
On GPT-5.5 and GPT-5.5 Pro, the system places implicit breakpoints at 2,048-token intervals, and retention is set with prompt_cache_retention, limited to 24h. Earlier models such as GPT 5.4, GPT 5.1, GPT 5 and GPT 4.1 support both in_memory retention, which typically lasts around 5 to 10 minutes of inactivity, and the extended 24h option.
The cached-input rate depends on the model, but there is no extra write charge on these versions. That makes caching free to try. The worst case is that you pay the normal price.
GPT-5.6 and the Write Fee
GPT-5.6 and later changed the rules, and the new setup is closer to Claude's:
Two modes: with prompt_cache_options.mode set to implicit, a breakpoint lands at the end of the latest eligible message. With explicit, you mark each breakpoint yourself using prompt_cache_breakpoint.
Pricing: cache reads cost 0.1x the uncached input rate, and cache writes cost 1.25x.
Lifetime:prompt_cache_options.ttl accepts one value, 30m, which is also the default.
Minimum: 1,024 visible input tokens.
Routing: a cache routing parameter on the request lets you keep separate cache accounting per customer, user or workspace.
{
"model": "YOUR_GPT_5_6_MODEL",
"prompt_cache_options": { "mode": "explicit" },
"input": [
{
"role": "developer",
"content": [{
"type": "input_text",
"text": "Stable rubric and instructions...",
"prompt_cache_breakpoint": { "mode": "explicit" }
}]
},
{ "role": "user", "content": "The changing part of the request" }
]
}
The three current tiers on the platform are GPT 5.6 Terra, GPT 5.6 Luna and GPT 5.6 Sol. Check the live pricing table for each tier before you budget, because OpenAI lists some tiers with deeper cached-read rates.
Gemini: Two Caches in One
Google offers two separate mechanisms, and they bill differently. One happens on its own. The other is an object you create and manage.
Implicit Caching by Default
Implicit caching is enabled by default for all Gemini 2.5 and newer models. You change no code. If a request shares a common prefix with an earlier one, it is eligible for a hit, and the savings are passed on automatically.
The minimums are 2,048 tokens for Gemini 2.5 Flash and 2.5 Pro, and 4,096 tokens for Gemini 3.5 Flash, its newer Flash siblings and Gemini 3.1 Pro. Google's own tips match the prefix rule exactly: put large, common content at the beginning, and send requests with a similar prefix close together in time.
There is no promise on any single request, though. Implicit caching is opportunistic, and that is why some teams move to the explicit version.
Explicit Caches and Storage Fees
With explicit caching you create a cache object, then point requests at it.
The default lifetime is one hour, and you can set your own with ttl, for example "300s". You can later change the ttl or expire_time, and nothing else about the cache. Billing has three parts: reused tokens at a reduced rate, storage per token-hour for as long as the cache exists, and the normal price for everything outside the cache. Explicit caches have a minimum of 2,048 tokens on 2.5 models and 4,096 on the 3.x line. The Interactions API only supports the implicit kind.
Google has published reuse discounts between 75% and 90% depending on the model generation, so confirm the number for your model on the live pricing page.
Side by Side Numbers
Feature
Claude
OpenAI (GPT-5.6 and later)
Gemini
How you enable it
cache_control breakpoints, or one top-level field for automatic mode
Implicit by default, or explicit breakpoints
Implicit by default, plus explicit cache objects
Lifetime
5 minutes, or 1 hour on request
30 minutes
Implicit is not guaranteed. Explicit defaults to 1 hour
Write cost
1.25x for 5 minutes, 2x for 1 hour
1.25x
Normal input price for implicit. Storage per token-hour for explicit
Read cost
0.1x on most models, 0.05x on Opus 5.5 and Sonnet 5.5
0.1x
Reduced rate, set per model
Minimum prefix
512 to 4,096 tokens by model
1,024 visible tokens
2,048 on 2.5, 4,096 on newer models
Control
Up to 4 breakpoints
Implicit or explicit mode
Cache object with a TTL you set
Which One Costs Less?
Use Claude's numbers to check the break-even point. A single 5-minute write at 1.25x plus one read at 0.1x costs 1.35x. Two uncached calls cost 2x. So one reuse already pays for the write. The 1-hour option writes at 2x, so you need two reads before it beats going uncached.
Now scale it up. Take a 20,000-token system prompt sent 1,000 times a day at the Opus 5.5 base rate of $4 per million tokens:
No caching: 20 million tokens at $4 per million is $80.00.
Caching with 50 cold restarts a day: 50 writes at $5 per million cost $5.00, and 950 reads at $0.20 per million cost $3.80, for $8.80 in total.
The user messages and the model output are billed the same either way, so this isolates the saving on the prefix. On OpenAI's older models with no write fee, the math is even simpler. On Gemini, the explicit route adds storage hours to the bill, so a cache that sits idle for most of the day can cost more than it saves.
Mistakes That Kill Cache Hits
Most failed caching setups are not provider problems. They come from three habits.
Timestamps in the Prefix
The classic bug is a line like "Current time: 14:32:07" near the top of the system prompt. The prefix differs on every request, so it can never match. Anthropic documents the same trap with a breakpoint placed on a block that holds a timestamp and a user message: the cache never hits, because no entry was written at any earlier position. The fix is to move the breakpoint to the last block that stays identical across requests and put the changing line after it.
Reordering Tools and Messages
Order is part of the fingerprint. Shuffling tool definitions, sorting a list of retrieved documents differently, or serializing JSON with a new field order all produce a new prefix. With Claude, a change at one level invalidates that level and everything after it: edit a tool definition and the tools, system and messages caches all go. Adding or removing images invalidates the messages cache. Keep tool lists in a fixed order and serialize them the same way every time.
Cold Traffic Between Bursts
A 5-minute lifetime does not help a job that sends one request every ten minutes. Each call pays the write price and never sees a read. In that case you have three options: move to the 1-hour lifetime on Claude, send a cheap keep-alive request before the timer runs out, or batch the work so requests arrive close together.
💡 Rule of thumb: match the cache lifetime to the gap between requests, not to the length of the session.
Measuring Hits in Production
Do not trust a setup until the response says it worked. Every provider reports cache activity in the usage block of the response.
Fields to Log
Provider
Where to look
Claude
usage.cache_creation_input_tokens and usage.cache_read_input_tokens
OpenAI
usage.input_tokens_details.cached_tokens and, on GPT-5.6 and later, cache_write_tokens
Gemini
The cached token count in usage_metadata
On Claude, the input_tokens field counts only the tokens after the last breakpoint, not the whole prompt. The real total is cache_read_input_tokens + cache_creation_input_tokens + input_tokens, so compute your hit rate from that sum.
Log three numbers per request: cached tokens read, tokens written, and tokens billed in full. Then watch two signals. A read count that stays at zero after the second identical request means the prefix is changing or sits under the minimum. A write count that climbs on every request means the lifetime expires before the next call arrives.
Create Your Own Images with Picasso IA
Draft Prompts in Three Models
Caching switches live in each provider's own API, so the place to tune them is your code. Picasso IA is useful one step earlier, when you are deciding what the stable prefix should say:
Open Claude Sonnet 5, paste your long system prompt, and ask it to tighten the wording without changing the rules.
Compare the three answers, keep the prompt that behaves well everywhere, and freeze it as your cached prefix.
A frozen prompt is also a stable one, which is exactly what a cache wants.
The photographs in this article were all made with P Image, each one from a single descriptive prompt about lens, light and texture. You can do the same in a few minutes. Try Seedream 4.5 for sharp 4K detail, Flux 2 Pro for photographic realism, Imagen 4 for natural scenes, or Nano Banana Pro for polished results.
Write a prompt that names the subject, the light and the lens, run it, then change one detail and run it again. The fastest way to get a feel for a model is to start experimenting with your own images on Picasso IA today.