Large Language ModelsGenerate imagesGenerate videos

Prompt Caching Explained: Claude, OpenAI and Gemini Compared

Prompt caching can cut the cost of repeated input to a fraction, but Claude, OpenAI and Gemini each run it differently. See how breakpoints, lifetimes, write fees and minimum sizes compare, with a worked cost example and the mistakes that stop cache hits.

Prompt Caching Explained: Claude, OpenAI and Gemini Compared
Cristian Da Conceicao
Founder of Picasso IA

Every request you send to a large language model starts with the same expensive ritual. The system prompt, the tool definitions, the pasted documentation and the few-shot examples all get read again, token by token, at full price, even though nothing in them has changed since the last call. Prompt caching ends that waste. The provider stores the processed form of your repeated prefix and charges a fraction of the normal input price when the next request reuses it. Think of a restaurant's mise en place: the cook chops the shallots once before service, and every order after that is assembled from ready containers.

Claude, OpenAI and Gemini all offer the idea, but they disagree on almost everything else: who decides what gets cached, how long it lives, what it costs to write, and how big a prompt must be before it qualifies. This article lines the three up using numbers from their current documentation, so you can pick a setup and dodge the mistakes that quietly switch caching off.

What Prompt Caching Actually Does

When a model reads your prompt, it does real work for every token before it can write the first word of the answer. That work is identical every time the beginning of the prompt is identical. Prompt caching lets the provider keep the result of that work for a short window and reuse it, so the repeated part is processed once and billed at a discount afterward.

A librarian's hand pulling a drawer from an oak card catalog, the way a cached prefix is pulled instead of rebuilt

Picture a card catalog. The provider files the processed prefix under a fingerprint of its exact content. The next request that produces the same fingerprint pulls the drawer instead of rebuilding the contents from scratch.

The Prefix Rule

Caching works from the first token forward. The cached part has to be an exact, unbroken match for the start of your request, and the first difference ends the match. Everything after that point is processed and billed at the regular rate.

A long shelf of identical green volumes with one book pulled out of line

That single rule shapes every provider's advice. Put the content that never changes at the top, and push everything that changes (the user's question, the current date, retrieved snippets) to the bottom. A good top section usually holds:

  • Tool definitions that stay the same between calls
  • System instructions and style rules
  • Reference material such as a manual, a contract or a codebase summary
  • Few-shot examples you reuse on every request

💡 Quick check: if you printed two requests and highlighted what they share, the highlighted part must be one solid block that begins at the very first character. A shared paragraph in the middle does not count.

Where the Savings Come From

Three things improve when the prefix is reused:

  • Cost: cached tokens bill at a small fraction of the normal input rate, while only the new tail is charged in full.
  • Latency: a long prompt means a longer wait before the first output token. Skipping the reprocessing shortens that wait, and the gain grows with the length of the prefix.
  • Rate limits: Anthropic states that cache reads are not deducted from your rate limit, so cached traffic leaves room for more requests.

Agent loops, document question answering, long chat histories and classifiers with a big rubric benefit the most. A prompt that is different every time gets nothing.

Claude: Explicit Control

Claude gives you the most direct control of the three. You decide where the cacheable prefix ends, and you decide how long it should live.

Breakpoints and Automatic Mode

You mark a content block with cache_control of type ephemeral, and everything from the start of the request up to and including that block becomes the cached prefix. The order is fixed: tools first, then system, then messages. You can set up to four breakpoints, and the API returns a 400 error if you try to place a fifth block-level one.

There is also an automatic mode. Add a single cache_control field at the top level of the request, and the system applies the breakpoint to the last cacheable block, then moves it forward as the conversation grows.

{
  "model": "claude-opus-5-5",
  "cache_control": { "type": "ephemeral" },
  "system": "Long, stable instructions go here...",
  "messages": [
    { "role": "user", "content": "Today's question" }
  ]
}

A leather ledger with four ribbon bookmarks marking different pages

One detail catches long agent sessions. When the system looks for an earlier matching entry, it checks at most 20 positions back from each breakpoint. If a turn adds many blocks, the previous entry can fall outside that window, and the fix is a second breakpoint placed earlier in the prompt.

Prices and Lifetimes

The default lifetime is 5 minutes, and every hit refreshes the timer at no charge. If your traffic comes in slow bursts, you can opt into one hour with "ttl": "1h" at a higher write price.

ItemMultiplierOpus 5.5 example (per million tokens)
Base input1x$4.00
5-minute cache write1.25x$5.00
1-hour cache write2x$8.00
Cache read0.05x on this model$0.20

Most Claude models read from the cache at 0.1x of the base input price. The Opus 5.5 and Sonnet 5.5 tiers read at 0.05x, and some newer tiers go lower still.

Minimum Sizes by Model

The threshold depends on the model, and it has moved a lot between generations:

💡 Silent failure: a prompt shorter than the minimum is processed without caching, and no error is returned. The only sign is a cache field that stays at zero.

OpenAI: Breakpoints Arrive

OpenAI's story has two chapters. For a long time caching was fully automatic. With GPT-5.6 it gained breakpoints and a write fee, which makes it look much more like Claude.

Older Models: Fully Automatic

On GPT-5.5 and GPT-5.5 Pro, the system places implicit breakpoints at 2,048-token intervals, and retention is set with prompt_cache_retention, limited to 24h. Earlier models such as GPT 5.4, GPT 5.1, GPT 5 and GPT 4.1 support both in_memory retention, which typically lasts around 5 to 10 minutes of inactivity, and the extended 24h option.

A warehouse conveyor sorting identical parcels while a worker watches with a coffee

The cached-input rate depends on the model, but there is no extra write charge on these versions. That makes caching free to try. The worst case is that you pay the normal price.

GPT-5.6 and the Write Fee

GPT-5.6 and later changed the rules, and the new setup is closer to Claude's:

  • Two modes: with prompt_cache_options.mode set to implicit, a breakpoint lands at the end of the latest eligible message. With explicit, you mark each breakpoint yourself using prompt_cache_breakpoint.
  • Pricing: cache reads cost 0.1x the uncached input rate, and cache writes cost 1.25x.
  • Lifetime: prompt_cache_options.ttl accepts one value, 30m, which is also the default.
  • Minimum: 1,024 visible input tokens.
  • Routing: a cache routing parameter on the request lets you keep separate cache accounting per customer, user or workspace.
{
  "model": "YOUR_GPT_5_6_MODEL",
  "prompt_cache_options": { "mode": "explicit" },
  "input": [
    {
      "role": "developer",
      "content": [{
        "type": "input_text",
        "text": "Stable rubric and instructions...",
        "prompt_cache_breakpoint": { "mode": "explicit" }
      }]
    },
    { "role": "user", "content": "The changing part of the request" }
  ]
}

The three current tiers on the platform are GPT 5.6 Terra, GPT 5.6 Luna and GPT 5.6 Sol. Check the live pricing table for each tier before you budget, because OpenAI lists some tiers with deeper cached-read rates.

Gemini: Two Caches in One

Google offers two separate mechanisms, and they bill differently. One happens on its own. The other is an object you create and manage.

Implicit Caching by Default

Implicit caching is enabled by default for all Gemini 2.5 and newer models. You change no code. If a request shares a common prefix with an earlier one, it is eligible for a hit, and the savings are passed on automatically.

The minimums are 2,048 tokens for Gemini 2.5 Flash and 2.5 Pro, and 4,096 tokens for Gemini 3.5 Flash, its newer Flash siblings and Gemini 3.1 Pro. Google's own tips match the prefix rule exactly: put large, common content at the beginning, and send requests with a similar prefix close together in time.

There is no promise on any single request, though. Implicit caching is opportunistic, and that is why some teams move to the explicit version.

Explicit Caches and Storage Fees

With explicit caching you create a cache object, then point requests at it.

cache = client.caches.create(
    model="gemini-2.5-flash",
    config=types.CreateCachedContentConfig(
        system_instruction="Long, stable instructions...",
        contents=[big_document],
        ttl="3600s",
    ),
)

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Today's question",
    config=types.GenerateContentConfig(cached_content=cache.name),
)

Pallets stacked on cold-storage racking, a picture of paying for space by the hour

The default lifetime is one hour, and you can set your own with ttl, for example "300s". You can later change the ttl or expire_time, and nothing else about the cache. Billing has three parts: reused tokens at a reduced rate, storage per token-hour for as long as the cache exists, and the normal price for everything outside the cache. Explicit caches have a minimum of 2,048 tokens on 2.5 models and 4,096 on the 3.x line. The Interactions API only supports the implicit kind.

Google has published reuse discounts between 75% and 90% depending on the model generation, so confirm the number for your model on the live pricing page.

Side by Side Numbers

Three notebooks, a calculator and receipts laid out on an oak desk for a cost comparison

FeatureClaudeOpenAI (GPT-5.6 and later)Gemini
How you enable itcache_control breakpoints, or one top-level field for automatic modeImplicit by default, or explicit breakpointsImplicit by default, plus explicit cache objects
Lifetime5 minutes, or 1 hour on request30 minutesImplicit is not guaranteed. Explicit defaults to 1 hour
Write cost1.25x for 5 minutes, 2x for 1 hour1.25xNormal input price for implicit. Storage per token-hour for explicit
Read cost0.1x on most models, 0.05x on Opus 5.5 and Sonnet 5.50.1xReduced rate, set per model
Minimum prefix512 to 4,096 tokens by model1,024 visible tokens2,048 on 2.5, 4,096 on newer models
ControlUp to 4 breakpointsImplicit or explicit modeCache object with a TTL you set

Which One Costs Less?

Use Claude's numbers to check the break-even point. A single 5-minute write at 1.25x plus one read at 0.1x costs 1.35x. Two uncached calls cost 2x. So one reuse already pays for the write. The 1-hour option writes at 2x, so you need two reads before it beats going uncached.

Now scale it up. Take a 20,000-token system prompt sent 1,000 times a day at the Opus 5.5 base rate of $4 per million tokens:

  • No caching: 20 million tokens at $4 per million is $80.00.
  • Caching with 50 cold restarts a day: 50 writes at $5 per million cost $5.00, and 950 reads at $0.20 per million cost $3.80, for $8.80 in total.

The user messages and the model output are billed the same either way, so this isolates the saving on the prefix. On OpenAI's older models with no write fee, the math is even simpler. On Gemini, the explicit route adds storage hours to the bill, so a cache that sits idle for most of the day can cost more than it saves.

Mistakes That Kill Cache Hits

Most failed caching setups are not provider problems. They come from three habits.

A hand stamping a different date on each of a stack of identical forms

Timestamps in the Prefix

The classic bug is a line like "Current time: 14:32:07" near the top of the system prompt. The prefix differs on every request, so it can never match. Anthropic documents the same trap with a breakpoint placed on a block that holds a timestamp and a user message: the cache never hits, because no entry was written at any earlier position. The fix is to move the breakpoint to the last block that stays identical across requests and put the changing line after it.

Reordering Tools and Messages

Order is part of the fingerprint. Shuffling tool definitions, sorting a list of retrieved documents differently, or serializing JSON with a new field order all produce a new prefix. With Claude, a change at one level invalidates that level and everything after it: edit a tool definition and the tools, system and messages caches all go. Adding or removing images invalidates the messages cache. Keep tool lists in a fixed order and serialize them the same way every time.

Cold Traffic Between Bursts

A 5-minute lifetime does not help a job that sends one request every ten minutes. Each call pays the write price and never sees a read. In that case you have three options: move to the 1-hour lifetime on Claude, send a cheap keep-alive request before the timer runs out, or batch the work so requests arrive close together.

💡 Rule of thumb: match the cache lifetime to the gap between requests, not to the length of the session.

Measuring Hits in Production

Do not trust a setup until the response says it worked. Every provider reports cache activity in the usage block of the response.

A hand holding a silver stopwatch at the finish line of a running track

Fields to Log

ProviderWhere to look
Claudeusage.cache_creation_input_tokens and usage.cache_read_input_tokens
OpenAIusage.input_tokens_details.cached_tokens and, on GPT-5.6 and later, cache_write_tokens
GeminiThe cached token count in usage_metadata

On Claude, the input_tokens field counts only the tokens after the last breakpoint, not the whole prompt. The real total is cache_read_input_tokens + cache_creation_input_tokens + input_tokens, so compute your hit rate from that sum.

Log three numbers per request: cached tokens read, tokens written, and tokens billed in full. Then watch two signals. A read count that stays at zero after the second identical request means the prefix is changing or sits under the minimum. A write count that climbs on every request means the lifetime expires before the next call arrives.

Create Your Own Images with Picasso IA

Draft Prompts in Three Models

Caching switches live in each provider's own API, so the place to tune them is your code. Picasso IA is useful one step earlier, when you are deciding what the stable prefix should say:

  1. Open Claude Sonnet 5, paste your long system prompt, and ask it to tighten the wording without changing the rules.
  2. Run the same task through GPT 5.6 Terra and Gemini 3.5 Flash.
  3. Compare the three answers, keep the prompt that behaves well everywhere, and freeze it as your cached prefix.

A frozen prompt is also a stable one, which is exactly what a cache wants.

A photographer arranging printed photos across a table in a sunlit loft

The photographs in this article were all made with P Image, each one from a single descriptive prompt about lens, light and texture. You can do the same in a few minutes. Try Seedream 4.5 for sharp 4K detail, Flux 2 Pro for photographic realism, Imagen 4 for natural scenes, or Nano Banana Pro for polished results.

Write a prompt that names the subject, the light and the lens, run it, then change one detail and run it again. The fastest way to get a feel for a model is to start experimenting with your own images on Picasso IA today.

Share this article