Lesson 4 of 6
Reuse stable prefixes; verify cache hits
Use prompt caching as a provider-specific billing and latency tool, keep prefixes stable, and confirm cache reads in usage — without assuming a universal discount or a smaller request.
Four outcomes, only one of which is “fewer tokens”
Token reduction means you sent or generated fewer tokens. Billing discount means the same tokens were priced cheaper (cache reads, batch APIs). Latency includes time to first token and total time to finish; measure them separately. Quality is whether the answer still meets the check.
Prompt caching is usually a discount + latency story. Official docs for OpenAI, Anthropic, and xAI all state that caching does not change how output is generated, and cached prefixes still occupy the context window. Anthropic: caching “changes what you pay for those tokens, not whether they count.” xAI: cached prompt tokens still count toward TPM, and long-context price tiers use total prompt tokens including cached tokens. OpenAI: cached prompts still count toward rate limits. Do not report a cache hit as “we reduced input_tokens to zero.”
Prefix stability is the whole trick
Caches match an exact prefix. Put stable material first: tools/schemas, system or developer instructions, reference docs. Put the changing user question last. Do not inject timestamps, shuffled tool lists, or rewritten history into the prefix.
Provider specifics (re-read current docs; they change):
- Anthropic:
cache_control(automatic or explicit). Hierarchy istools→system→messages. Writes cost more than base input; reads cost less.input_tokensexcludes cached spans. Minimum cacheable length is model-specific. Hits need identical bytes; TTL is documented (default 5 minutes, optional 1 hour). - OpenAI: prefix caching; older models implicit, GPT-5.6+ families add explicit breakpoints and a cache-write rate.
cached_tokensis a subset of input. Hits are not guaranteed (routing, TTL, minimum length). - xAI: automatic from the start of
messages. Setx-grok-conv-id(Chat) orprompt_cache_key(Responses). Do not edit earlier messages. Reasoning models: omittingreasoning_contentis a common miss. Hits are not guaranteed (eviction, routing).
Read the cache fields; do not guess a percentage
Verify on the second identical-prefix request, inside the documented lifetime:
| Provider | Miss / write | Hit |
|---|---|---|
| Anthropic | cache_creation_input_tokens > 0 |
cache_read_input_tokens > 0 |
| OpenAI / xAI | cached_tokens = 0 (first request or eviction) |
cached_tokens > 0 |
If both Anthropic cache fields stay 0, the prefix may be under the minimum — there is often no error. OpenAI GPT-5.6+ may report cache_write_tokens billed above the uncached input rate; later reads are cheaper. There is no course-wide savings percentage. Published “up to 90%” figures are marketing ceilings on cache-read rates, not a promise for your trace. Look up today’s pricing page for your model. Cache writes can make a one-off request cost more.
A worked comparison (labelled simulated)
Run two requests that share a long static prefix and differ only in the last user line. Log occupancy, cached read, cache write, output, and clock time if you have it.
Example / simulated — not your result: Request 1 occupancyIn 12 800, cached_read 0, cache_write 12 400, output 180. Request 2 occupancyIn 12 850, cached_read 12 400, cache_write 0, output 160. Occupancy did not collapse; the second request still held ~12.8k input tokens. What changed was the category mix (and maybe TTFT). If request 2 shows cached_read 0, debug prefix drift (tool order, timestamp, dropped reasoning block) before you claim caching “doesn’t work.”