Prompt caching does not reduce anything by itself. It reuses work when the rendered prefix of a request matches a prefix the provider already processed, and it fails silently when it does not. So the practice comes down to three questions: what counts as the prefix, what breaks the match, and how do you verify that a hit actually happened. This guide answers all three, using the two mechanisms you are most likely to meet in production, and keeps the mechanics tied to each provider's own documentation rather than folklore. Used well, it is also one of the few levers that reduce LLM API cost without changing what your product does, but it is the second thing you fix, never the first.
Prompt caching, two mechanisms, one rule
OpenAI caches automatically. Per the provider's prompt caching guide, the cache reuses the rendered prefix of your request, which can include text, images, documents, and supported audio, and reuse requires the entire rendered prefix to match. The prefix is not just your first message: it is assembled from the hidden system content, tool definitions and schemas, developer instructions, and the conversation history, roughly in that order. If content or a relevant setting changes before a breakpoint, everything after that change cannot match the existing entry. There is a minimum cacheable length, and it varies by model, so the current threshold belongs to the official page, not to your memory.
Anthropic works by explicit marks. Its prompt caching documentation describes two modes: automatic caching, where a single top-level cache_control field places the breakpoint on the last cacheable block and moves it forward as the conversation grows, and explicit cache_control breakpoints on individual content blocks when you want fine-grained control. Two mechanics matter for practice. Writes happen only at your breakpoint, and the hash is cumulative, so changing any block at or before the breakpoint produces a different hash on the next request. Reads look backward from your breakpoint, checking earlier positions for an entry that a prior request already wrote. The asymmetry is the point: a request writes at most one cache entry per breakpoint, and every later request that matches that hash reuses it, while everything after the mark is processed fresh. Get the mark right, and the reads follow.
Different plumbing, one shared rule: the stable parts of a request must sit in front, and everything volatile must sit behind them. Cache lifetime options and cacheable-length thresholds are documented per provider and change over time, so read the current values on each provider's page rather than copying them from anywhere, including from this article. If you consume models through a compatible platform such as String AI, the platform's documentation is the contract for which behaviors carry over on its endpoint.
Prompt cache not working? The silent breakers
A cache that misses does not throw an error. It just processes everything again, and your only clue is a field in the usage data. These are the writers that break prefixes most often, each with the fix.
A timestamp, UUID, or random seed in the prefix. Anything that changes per request poisons every token after it. Symptom: hits near zero despite identical-looking prompts. Fix: move volatile values after the breakpoint, or drop them entirely.
User input assembled ahead of the stable instructions. Concatenating the user's question before your system rules means the prefix differs from the first volatile token onward. Symptom: the system prompt looks cached, and is not. Fix: freeze the front of the template, and append dynamic content last.
Tool definitions edited per request. OpenAI's guidance is direct that changes to tool names, descriptions, schemas, or ordering change the prefix, and tools sit near the front of the stack. Symptom: cache misses that started the day someone "cleaned up" the tool list. Fix: version your tool block as a unit, and stop reordering it casually.
Unstable multimodal ordering. Images and documents are part of the rendered prefix too, so the same assets attached in a different order are a different prefix. Symptom: misses that correlate with a UI that does not preserve attachment order. Fix: sort attachments deterministically before sending.
A context block that mutates every turn. Rolling summaries, running notes, or "current state" blobs that get rewritten each call will break the prefix they touch, even when the surrounding conversation is untouched. Symptom: multi-turn sessions that never hit after the first exchange. Fix: if it must change, keep it after the breakpoint, and cache only what truly persists.
A breakpoint placed on the wrong side of the churn. Anthropic's documentation describes the failure mode precisely: with the breakpoint on a block that changes every request, you pay for a fresh write every time and never read a hit. Symptom: cache writes with no reads. Fix: move the breakpoint to the last block that stays stable across requests.
One habit contains all six fixes: treat the prefix as a template with a frozen front section, and give that section a version you can see in your own logs. Keep a short note next to the template describing which sections may change, which are frozen, and where the breakpoints sit, so the next engineer does not have to rediscover the same rules by watching hit rates collapse.
Verifying cache hits in your usage data
Reading cached tokens usage is the difference between believing and knowing. The verification loop is small, and it is worth running in this order.
- Read the fields. In OpenAI responses, cached input shows up in the usage object's token details, alongside cache write accounting when a prefix is first stored. Anthropic reports cache writes and cache reads as distinct events, which is convenient, because a growing read count is the signature you want.
- Repeat the same prefix twice. Identical prefix, two calls, compare the usage objects. The second call is where a hit must appear; if both look the same, the prefix is not actually identical. If the fields move but never accumulate, re-check the breakpoint placement from the previous section before changing anything else.
- Run a minimal A/B. Keep the prefix untouched in one variant, insert a single variable in the other, and compare. Start with one small call in a terminal before wiring code, the same discipline as the quickstart, because a minimal reproduction beats a full application when you are testing mechanics.
- Log a fingerprint. Hash the prefix you send, and log it next to the cache fields. The moment hit rates change, the diff of two fingerprints tells you what changed, which beats re-reading code archaeology.
- Audit on a schedule. Periodically confirm against your usage records that reads are actually happening, rather than trusting a dashboard you set up months ago and never touched. And keep in mind that usage fields and response shapes differ across endpoints, which is exactly the variance the compatibility checklist exists to pin down.
One extra tool if you are on OpenAI: a cache key can separate cache accounting per user or workflow, which makes verification and per-customer explanation cleaner.
When caching is not worth it
Caching rewards workloads with repetitive prefix structure, and it is honest to say that not every workload qualifies.
- The prefix is genuinely unique per request. Long-tail personalized prompts where everything before the breakpoint differs have nothing to reuse, and no amount of cache tuning changes that.
- You call once per artifact. Single-shot generation with no follow-up has no second request to hit the cache.
- The prompt is short. Below the provider's minimum cacheable length, there is no eligible cache entry at all, and the threshold varies by model, so check the official pages rather than assuming.
- You have not measured hit rate. Adding a cache layer to a workload you never measured is how teams add complexity and call it optimization. Measure first: if your stable prefix is not actually stable in production logs, the fix is prefix discipline, not a caching project.
- Your client works against you. Caching depends on what the client actually transmits; tooling that injects session identifiers, reorders attachments, or rewrites instructions between requests undermines the prefix regardless of how good your template is. Fix the client first, or caching will be a permanent mystery.
Where to go next
- See the spend before tuning it. Usage and cost attribution is the measurement layer this article assumes.
- Triage what breaks. When caching behaves oddly in ways that look like errors, the error code guide separates the layers, including the ones that never deserve a retry.
- Where prefixes get shaped. Client and tool configuration is upstream of everything here, and the Cursor endpoint walkthrough is a concrete example of that layer.
Fix the prefix first, verify with your own usage data, and only then talk about what caching saved.