Claude prompt caching: what it costs and why it stops working
Prompt caching is the largest cost lever most Claude applications have, and the easiest to break without noticing. A cache that stops matching raises no error. It shows up later, as a bill. This guide explains the mechanism first, then the money, then the checks that tell you whether it is working.
In short
- Reading cached input costs about a tenth of the normal price. Writing it costs 1.25 times the normal price for five minutes, or 2 times for one hour.
- The cache matches an exact prefix. One changed byte early in the request makes everything after it miss.
- A broken cache raises no error, and it costs 25% more than no caching at all.
- Choose the lifetime by the gap between requests: under five minutes, keep the default. Between five and sixty minutes, use one hour.
- The usage block in every response tells you whether caching works. Check it after every change to how prompts are built.
What prompt caching is
Claude reads your whole prompt on every request: the tool definitions, the system prompt, the conversation so far and the new message. In most applications the first three barely change from one request to the next, and you pay full price for them every time.
Prompt caching lets Anthropic keep the work it already did on the start of your prompt. You mark where the reusable part ends. That mark is called a breakpoint. The next request that begins with exactly the same content reads it from the cache at a fraction of the price.
The third row is the whole risk in one picture. Nothing failed. The request succeeded and the answer was fine. The cached part simply cost over twelve times as much as it did one request earlier.
What each token costs
With caching switched on, an input token has one of four prices, all expressed as a multiple of the model's normal input price.

| Model | Input | Write, 5 min | Write, 1 hour | Read |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $12.50 | $20 | $0.25 |
| Claude Opus 5.5 | $4 | $5 | $8 | $0.20 |
| Claude Sonnet 5.5 | $2 | $2.50 | $4 | $0.20 |
| Claude Haiku 4.5 | $1 | $1.25 | $2 | $0.10 |
One detail in this table is easy to miss. Reads are a tenth of the input price on most models, but only a twentieth on Claude Opus 5.5 and a fortieth on Claude Fable 5.1. The more expensive the model, the more a miss costs you compared with a hit.
When it pays for itself
A cache write is a small bet that the same prefix will be used again. With the five-minute lifetime the bet pays off on the second request: one write and one read cost 1.35 times the input price, against 2 times without caching. With the one-hour lifetime you need three requests.

There is one more condition. A prefix must reach a minimum length before it can be cached at all.
| Minimum prefix | Models |
|---|---|
| 512 tokens | Claude Fable 5.1, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5 |
| 1,024 tokens | Claude Sonnet 5, Claude Opus 4.8 |
| 4,096 tokens | Claude Haiku 4.5 |
The prefix rule
Every request is assembled in a fixed order: tool definitions, then the system prompt, then the messages. The cache matches from the first byte up to a breakpoint. If anything before the breakpoint differs, the match fails from that point on.
That gives one design rule. Put what never changes first, and what changes on every request last.
const response = await client.messages.create({
model: 'claude-sonnet-5-5',
max_tokens: 1024,
// Automatic caching: the breakpoint follows the end of the conversation
cache_control: { type: 'ephemeral' },
system: [
// A fixed breakpoint after the part that never changes
{ type: 'text', text: SYSTEM_PROMPT, cache_control: { type: 'ephemeral' } },
],
messages,
})This is the combination that holds up best in production. The breakpoint on the system prompt guarantees a read point for the large, stable part. The top-level setting moves a second breakpoint forward as the conversation grows, so each turn reads everything before it.
- A request can carry up to four breakpoints. Automatic caching uses one of them.
- A read can only land where an earlier request wrote. The cache does not find stable content on its own. It finds entries written at a breakpoint.
- From each breakpoint the system looks back at most 20 content blocks for an earlier entry. A single turn that adds more than that, such as a long run of tool calls, can miss the previous entry. Add a breakpoint partway through long turns.
- A cache entry becomes readable once the first response starts streaming. Ten identical requests fired at the same instant all pay the write price. Send one, wait for its first token, then send the rest.
What silently breaks the cache
None of these raise an error. Each one turns reads into writes.
| Cause | What happens | Fix |
|---|---|---|
| A timestamp, date or request ID in the system prompt | Every request has a new prefix | Move it to the latest user message |
| JSON written with keys in a changing order | The bytes differ although the data is the same | Serialize with sorted keys |
| Tools added, removed or reordered | Tools come first, so the whole request misses | Keep one fixed, sorted tool list |
| Switching models mid-conversation | Each model has its own cache | Keep a conversation on one model |
| Changing thinking or effort settings between requests | The conversation cache is invalidated | Fix the settings per route |
| A breakpoint placed after the unique part of the prompt | Every request writes an entry nobody reads | Put the breakpoint at the end of the shared part |
| A prefix under the model's minimum length | Nothing is cached | Check the minimum when you change models or shorten prompts |
| The same prompt sent from several workspaces | On the Anthropic API each workspace has its own cache | Send shared traffic from one workspace |
Some changes are safer than they look. Changing tool_choice from one request to the next keeps the tools and system prompt cached. Adding a message never affects what came before it.
Five minutes or one hour
A cache entry lives for five minutes by default, or one hour if you ask for it. Every read resets the clock at no charge. So the question is not how long your sessions last. It is how long the gaps are between requests that share a prefix.

- The clock runs from the start of a request, not from the end. If a response takes four minutes to generate, the next request has about one minute to start before a five-minute entry expires.
- You can mix lifetimes in one request, with one constraint: the one-hour breakpoint must come before any five-minute breakpoint.
- To request one hour, write cache_control: { type: "ephemeral", ttl: "1h" }.
- A request with max_tokens: 0 writes or refreshes the cache without generating a reply. It is useful at start-up, when the first real user should not pay the cold-start delay.
How to check that it works
Every response includes a usage block. It is the only reliable evidence that caching works.
"usage": {
"cache_read_input_tokens": 14000, // read from the cache, cheap
"cache_creation_input_tokens": 600, // written to the cache, at the write price
"input_tokens": 45, // after the last breakpoint, at full price
"output_tokens": 380
}The three input fields add up to the size of your prompt. If input_tokens looks surprisingly small, the rest was served from the cache.

- Healthy: cache_read_input_tokens covers almost the whole earlier conversation and grows each turn. cache_creation_input_tokens is about the size of the latest turn.
- Broken: cache_creation_input_tokens is close to the full prompt on every request and reads stay near zero.
- Not caching at all: both cache fields are zero. The prefix is probably under the minimum length, or there is no breakpoint.
To find the cause, log the full body of two consecutive requests and compare them. The first place they differ, before the newest message, is what breaks the cache.
A worked example
Take a 14,000-token system prompt sent on 100,000 calls a month, on a model priced at $2 per million input tokens. The table shows what that prefix alone costs.
| Situation | How it is billed | Per month |
|---|---|---|
| No caching | 1.4 billion tokens at $2 | $2,800 |
| Working cache | 98,000 reads at $0.20 and 2,000 cold writes at $2.50 | about $345 |
| Broken cache | 1.4 billion tokens written at $2.50 | $3,500 |
A working cache saves about $2,450 a month on this one prompt. A broken cache costs $700 a month more than doing nothing, which is why the usage check matters more than the set-up.
Caching on Amazon Bedrock
If you run Claude through Amazon Bedrock to use AWS credits, caching is available there too, with the same minimum lengths and both lifetimes on current models. A few things differ.
- Caches are separated per organization, not per workspace.
- AWS notes that cross-region inference can cause more cache writes at times of high demand, because a request may be served from a region that has not seen the prefix.
- Caching works with on-demand inference only, not with Bedrock batch inference.
- Claude Opus 4.6 and earlier models use the older Bedrock APIs, where the usage fields are named cacheReadInputTokens and cacheWriteInputTokens, and where automatic caching is not available.
- Amazon CloudWatch publishes CacheReadInputTokenCount and CacheWriteInputTokenCount for the Bedrock runtime. If your endpoint does not publish them, log the usage block from each response yourself.
Common questions
Does caching change the answers?
No. Caching changes what the input costs and how fast the response starts. The model receives the same prompt.
Can other customers read my cached prompts?
No. Caches are never shared between organizations. On the Anthropic API they are also separated per workspace.
Does caching make responses faster?
Yes, for long prompts. The cached part does not have to be processed again, so the first token arrives sooner.
What can be cached?
Tool definitions, system prompts, text, images, documents, tool calls and tool results. Thinking blocks cannot be marked directly, but they are cached as part of earlier turns.
Do cached tokens count toward rate limits?
On most models, cache reads are not counted against the input tokens per minute limit, so a working cache also raises your throughput.
Should I cache a prompt that changes on every request?
No. If the start of the prompt is different each time, there is nothing to reuse, and you pay the write premium for nothing.