Prompt caching lets the model reuse a chunk of your prompt it has already processed, billing it at a fraction of the normal input rate. For anything with a large stable prefix — a long system prompt, a document, a growing conversation — it's the single biggest saving available and it costs you no quality at all.

The multipliers

Relative to the model's base input price:

  • Writing to a 5-minute cache costs 1.25x
  • Writing to a 1-hour cache costs 2x
  • Reading from cache costs 0.1x on most models

Claude Fable 5.1 is the exception and it's a big one: cache reads there cost 0.025x, which is $0.25 per million tokens against a $10 base rate. On the most expensive model, cached input is the cheapest input anywhere.

The break-even

Simple arithmetic. A 5-minute cache write costs 1.25x and each read costs 0.1x, so you're ahead after a single read. A 1-hour write costs 2x, so you need two reads.

That's a low bar. Almost any repeated prompt clears it.

What to cache

Anything that doesn't change between requests, in the order the request is assembled: tool definitions first, then the system prompt, then the stable early part of the conversation.

What must come after the cached part is anything that varies: the current question, timestamps, per-request identifiers.

The thing that silently breaks it

Caching is a prefix match. Any byte that changes anywhere in the cached prefix invalidates everything after it. The classic mistakes are a current timestamp in the system prompt, a tool list assembled in a different order each time, or JSON serialised without stable key ordering.

The symptom is that you're paying the 1.25x write price over and over and never getting the 0.1x read.

cache-audit.txt
My cache hit rate is low or zero. Here's how I build my requests:
<paste the system prompt assembly, tool definitions, and message
construction>

Find the silent invalidators — anything in the cached prefix that
changes between requests. Check specifically for timestamps, unordered
JSON, tool lists built in varying order, and per-request IDs placed
before the cache breakpoint.

Then show me the corrected assembly, and tell me which usage field I
should watch to confirm it worked.

The field to watch is the cache read token count in the response. If it's zero across repeated requests, something is still invalidating.

Where it doesn't apply

If you're vibe coding through a chat interface or a coding tool, caching is handled for you and there's nothing to configure. This matters when you're building your own product on the API.

Caching is a free win: same output, lower bill. Do it before you touch anything that trades quality for cost.