The cheapest token is the one you did not send
Evaluate lookup, prompt caching and answer reuse against the actual workload. Stable prompt structure can reduce processing cost, while answer caches need explicit access, freshness and invalidation rules.
Keep development and test groups separate before comparing performance. No measured results are shown.
Cost reduction should begin with the work the service performs. Some requests may be handled through a lookup, reusable prompt content or a retained answer. Others need a fresh model call. Compare these options before assuming that a smaller model is the best response to a rising bill.
Prompt structure affects provider-side caching, while answer reuse depends on data, permissions and freshness. Both deserve a design decision. They can be introduced later, but a change to prompt order or answer selection needs evaluation against the service’s acceptance conditions.
Treat caching as a scoped design option with measurable benefits and failure conditions. A repeated prefix can reduce processing cost where the provider supports it. The achievable saving depends on the request structure, cache rules and actual hit rate.
Four ways not to send a token
- Lever 01 Do not call the model A lookup, a rule, a regular expression or a stored answer. Use where a deterministic route meets the acceptance conditions.
- Lever 02 Reuse the prefix Provider-side prompt caching charges a reduced rate for a repeated prefix. It requires the prefix to actually repeat, byte for byte.
- Lever 03 Answer from a previous answer A semantic cache over prior questions and responses, with an explicit freshness rule and an explicit blast radius when it is wrong.
- Lever 04 Send a smaller context Fewer retrieved passages, shorter history, a trimmed tool schema. Evaluate missing evidence and task quality alongside cost.
Identify tasks that a deterministic route can meet reliably, such as an approved lookup. Measure their frequency rather than assuming a fixed share of traffic. A stored answer needs an owner, a source and a review rule even when no language model is called at request time.
How prefix structure affects reuse
Anthropic’s prompt caching documentation describes exact matching of cacheable prefixes across tools, system content and messages, subject to minimum lengths and cache rules. Stable content before variable content can improve reuse. Confirm the applicable model, pricing and lifetime rather than assuming every provider caches in the same way.
A candidate structure keeps repeated material before request-specific content:
- The system instruction for the selected version
- The tool and schema definitions for that release
- Policy or reference material that is stable across the eligible requests
- The retrieved passages for this request
- The conversation so far
- The user’s current message
Putting a changing value early in otherwise shared content can shorten the reusable prefix. It does not necessarily eliminate every cache hit. Inspect the provider’s usage records and compare observed cost and behaviour for the proposed structure.
Evaluate prompt-order changes as well as their cost. Reordering may affect the model’s output even when the supplied information is unchanged. Measure quality, latency and cache use together before accepting the revised template.
Govern the reused answer
Semantic caching reuses an answer for a sufficiently similar request. Similar wording does not establish equivalent meaning, entitlement or data freshness. The benefit and risk depend on how the service defines and tests those conditions.
Decide whose answer may be reused and how current it must be. A question-only cache can disclose an answer across users with different permissions. Include the relevant access scope and revalidate changes. Set freshness and invalidation rules against the underlying data: a policy answer may become obsolete immediately after revision, while an account answer may require a current query.
The third line is the one that saves an investigation later. If the record does not distinguish a generated answer from a served one, then the first question after an incident — did the model produce this, or did we hand back something from last Tuesday — has no answer.
What this does not tell you
Open-ended or novel work can have limited answer-reuse opportunities while still sharing prompt prefixes. Measure each mechanism separately. Relaxing semantic similarity to increase hits can return the wrong answer, so evaluate that change against meaning and consequence rather than hit rate alone.
Caching, smaller models and retrieval improvements can be complementary. Exact-prefix caching differs from reusing a generated answer, which can change what the user receives. Evaluate each option on operating cost and the service conditions it affects. The bill alone cannot validate an answer cache.
The architecture owner should document which material is stable, which answers may be reused and when reuse must stop. Ask the delivery team to measure the proposed savings and test access and invalidation. Accept the option that improves operating economics while preserving the required service behaviour.