Prompt and Semantic Caching - Complete Deep Dive
Prerequisites: Caching, Embeddings, Tokens and Cost Math Used in: Latency Engineering, GPU Cost Engineering, Deploying AI Features Build it: Lesson 21 - Latency and Cost Engineering implements this as runnable, tested code you can execute offline.
What is Caching in an LLM System?
Caching an LLM call means not paying for generation you have already paid for. The payoff is unusually large here, because unlike a database query a model call is expensive in three currencies at once β money, latency, and scarce accelerator capacity β so a cache hit returns a saving in all three.
The trap is that βLLM cachingβ names three completely different mechanisms that teams routinely discuss as one thing. They key on different material, they live in different places, and β the part that matters β they have different correctness properties. One of them cannot return a wrong answer. One of them changes only your bill. One of them can serve a confidently wrong answer to a question nobody has asked before. Treating them as interchangeable is how a cost optimisation becomes an incident.
Real-world analogy: think of a support desk. Exact-match caching is the agent recognising the identical ticket they closed an hour ago and pasting the same reply β safe, because it is the same question. Prompt caching is the agent already having the policy manual open on their desk instead of walking to the archive each time β same answer, less setup. Semantic caching is the agent deciding your question is βbasically the sameβ as one from last week and reusing that reply. Sometimes that is excellent service. Sometimes the difference between the two questions was the word βnotβ.
The Three Layers Side by Side
| Layer | What it keys on | Typical hit rate | Risk | Correctness property |
|---|---|---|---|---|
| Exact-match response cache | A hash of everything that determines the output - model id and version, sampling params, system prompt version, the full message list, tool schemas, retrieval context, tenant and user scope | Higher than teams expect on real traffic, because production traffic is far more repetitive than test traffic | Low, and entirely about key construction | Exact. Identical input gives the response the model already gave. A wrong answer here means your key was missing a component. |
| Provider-side prompt caching | A long stable prefix of the request, matched by the provider from the first token forward | Depends almost entirely on prompt ordering, not on query repetition | Low. It discounts repeated prefill, it does not change what is generated | Neutral. The model still runs. Output is unaffected, only the cost and time to first token change. |
| Semantic cache | An embedding of the query, matched by similarity against stored queries above a threshold | Can be high, and that is exactly the seduction | High. This is the only layer that can invent a wrong answer | Approximate. A hit asserts βsimilar enoughβ, which is not the same claim as βequivalentβ. The threshold is a correctness knob. |
Read the last column top to bottom. That progression β exact, neutral, approximate β is the whole page.
Layer 1 - Exact-Match Response Caching
Hash the request, store the response, return it on the next identical request. Unglamorous and the layer most teams under-build.
The single engineering decision is what goes into the key, and the rule is absolute: the key must contain every input that could change the output. Anything you leave out is something you are claiming does not matter.
def cache_key(req) -> str:
return sha256(canonical_json({
"model": req.model_id, # family and pinned snapshot, not just "the fast one"
"params": { # sampling changes the output distribution
"temperature": req.temperature,
"top_p": req.top_p,
"max_tokens": req.max_tokens,
"seed": req.seed,
},
"prompt_version": req.prompt_version, # a prompt edit must miss the cache
"messages": req.messages, # the full list, in order, verbatim
"tools": req.tool_schema_version, # different tools - different behaviour
"context_ids": sorted(req.retrieved_chunk_ids),
"context_version": req.index_version, # the corpus moved - the answer may move
"tenant": req.tenant_id, # isolation is part of the key
"locale": req.locale,
})).hexdigest()
The bug worth naming explicitly. Forget tenant and you will eventually serve one customerβs answer to another customer. That is simultaneously a correctness bug and a data-leak incident, and it is not hypothetical β it is the default outcome of keying on βthe question textβ alone, because two tenants asking βwhat is our current discount tierβ produce byte-identical questions and different correct answers. The same reasoning applies to user identity on anything personalised and to role on anything permission-scoped.
Two more things belong in the key that people forget because they are not part of the message list: the pinned model snapshot (a provider-side model update changes behaviour with no deploy on your side β see Deploying AI Features) and the retrieval index version (the same question over a re-indexed corpus is a different question).
Everything else is ordinary cache engineering, which Caching already covers: where it lives, eviction policy, memory budget. Use a shared store such as Redis or Memcached rather than a per-process dictionary, because per-process caches fragment your hit rate across replicas.
Layer 2 - Provider-Side Prompt Caching
Providers can cache the internal state of a long, stable prefix so that repeat requests sharing that prefix skip most of the prefill work. You do not manage the cache; you make your prompt eligible for it.
The consequence is sharper than it first sounds: prompt order becomes a cost decision. Prefix matching starts at the first token and stops at the first difference, so a single variable token near the top invalidates everything after it.
flowchart TD
subgraph Stable["Stable prefix - eligible for reuse"]
P1[System prompt at a pinned version]
P2[Tool schemas]
P3[Shared reference document or policy text]
P4[Few shot exemplars]
end
subgraph Variable["Variable tail - prefilled every time"]
V1[Retrieved chunks for this query]
V2[Conversation turn from this user]
end
P1 --> P2
P2 --> P3
P3 --> P4
P4 --> V1
V1 --> V2
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class P1,P2,P3,P4 data
class V1,V2 client
This is the opposite of what a naive template does. The instinctive template interpolates the userβs name, the current timestamp, the session id, or todayβs date into the opening line of the system prompt, then appends the long shared policy document underneath. That layout is prefix-hostile: the first line differs on every request, so nothing downstream can be reused. Moving the same variable material to the bottom changes no behaviour and makes the whole document cacheable.
Practical rules:
- Stable material first, variable material last. System prompt, tool schemas, shared documents, exemplars, then retrieved context, then the user turn.
- Pin and version the prefix. A prompt edit should be a deliberate, versioned change, not an accidental invalidation from a stray whitespace fix.
- Do not sort retrieved chunks non-deterministically. If ranking ties break randomly, an otherwise-identical request misses both this layer and Layer 1.
- Longer prefixes pay off more. The economics favour exactly the architectures that were already expensive: large system prompts, big tool catalogues, long shared context.
- Nothing about output changes. If the answer differs after enabling prefix caching, something else moved. Read the token accounting in Tokens and Cost Math before attributing it here.
Layer 3 - Semantic Caching, and Why It Is Dangerous
Embed the incoming query, search stored query embeddings, and if the nearest neighbour is similar enough, return its stored answer without calling the model. Mechanically it is a small retrieval system β see Embeddings.
Here is the core lesson of this page: similarity is not equivalence. Cosine similarity measures proximity in a representation trained to place related text near related text. It was never trained to certify that two questions have the same answer. A hit is your system asserting semantic equivalence on the strength of a number that does not mean equivalence.
The threshold is therefore not a tuning parameter, it is a correctness knob. Loosen it and the hit rate climbs while wrong answers start shipping β and they ship in the worst possible form, since a cached answer arrives fast, fluent, and previously validated.
The canonical failure is negation and small qualifiers, because the tokens that invert a required answer barely move an embedding:
| Stored query | Incoming query | Embedding distance | Correct answer |
|---|---|---|---|
| Is this expense reimbursable | Is this expense not reimbursable | Tiny | Inverted |
| Can I cancel within 30 days | Can I cancel after 30 days | Tiny | Opposite |
| What is the fee for domestic transfers | What is the fee for international transfers | Small | Different number |
| Does the policy cover this | Did the 2024 policy cover this | Small | Possibly different |
| How do I enable two factor auth | How do I disable two factor auth | Small | Different procedure |
Every row is a plausible support question and every row is a production incident if the threshold is loose.
Mitigations, in the order they matter:
- Tune the threshold on labelled pairs, and tune it conservatively. Build a set of query pairs hand-labelled as same-answer or different-answer, weighted toward the negation and qualifier cases above. Measure false-hit rate at each candidate threshold and pick for precision, not for hit rate. A semantic cache with a 20 percent hit rate and no wrong answers beats one with a 60 percent hit rate and occasional inversions.
- Scope the cache per tenant, and usually per user. A shared semantic cache across tenants compounds the approximation problem with a leak. Namespace the index.
- Never semantic-cache personalised or permission-dependent answers. If the correct response depends on who is asking or what they may see, similarity of the question tells you nothing about reusability of the answer.
- Exclude anything time-sensitive. Balances, statuses, availability, inventory, βlatestβ anything. The question is stable and the answer is not.
- Prefer narrow domains. Semantic caching works best on a bounded FAQ-shaped surface where paraphrase is common and stakes are low, and worst on open-ended reasoning.
- Verify on hit when the stakes justify it. A cheap model or a rule check can confirm the cached answer still addresses the new question β see Model Selection and Routing. You give back some of the saving to buy back correctness.
- Log every hit with its similarity score and both queries. Without this you cannot audit the layer at all, and you will not know what it got wrong.
The Fall-Through Path
The three layers compose. Cheapest and safest first, model last, write-back on the way out.
flowchart LR
U[Client request] --> K[Build key from full request identity]
K --> E[Exact match cache]
E -->|hit| R[Return stored response]
E -->|miss| S[Semantic cache lookup per tenant]
S -->|above threshold| R
S -->|below threshold or miss| M[Model call using cached stable prefix]
M --> R
M --> W[Write back]
W --> X[Exact cache entry with TTL]
W --> Y[Semantic index entry with query embedding]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class U client
class K,R edge
class E,S,X,Y data
class M service
class W async
Two details the diagram encodes deliberately. The write-back is asynchronous β a cache write must never sit on the response path, and a failed cache write must never fail the request. And write-back is selective: responses that failed validation, were refused, abstained, or errored do not belong in either cache. Caching a bad answer converts a transient failure into a durable one.
Invalidation, TTLs, and the Stampede
Invalidate by versioning the key, not by remembering to purge. Three things change the correct answer to an unchanged question: the corpus was re-indexed, the prompt was edited, or the model moved. Each one must appear in the key so the change produces a natural miss. The alternative β a purge step someone has to remember during a release β fails the first time a release is rushed. This is the same argument for release identity made in Deploying AI Features.
Version these into the key explicitly: prompt_version, model_id including the pinned snapshot, index_version, embedding_model_version for the semantic layer, and tool_schema_version. An embedding model change is especially unforgiving, because old stored vectors are not comparable with new ones β that cache must be rebuilt, not reused.
TTLs handle the freshness you cannot version. Version keys catch changes you know about; a TTL bounds your exposure to the ones you do not. Set it from how stale an answer may acceptably be: long for stable reference material, short for anything operational, zero for anything live.
The stampede. A popular key expires or goes cold and every concurrent request for it misses at once, all of them calling the model with the same input. On a GPU-bound path this is worse than a database stampede, because concurrency is the scarce resource β see GPU Cost Engineering. Standard mitigations transfer: a per-key lock so one request computes while the others wait on the result, serve-stale-while-revalidating for entries where slight staleness beats a latency spike, and staggered TTLs so entries written together do not expire together. Coalescing concurrent identical requests into one model call is the single highest-value version of this.
Measuring the Cache
Hit rate is the metric everyone reports and it is only half the instrumentation. Track per layer:
- Hit rate, per layer, segmented by route. An aggregate hit rate hides the fact that your cheap FAQ route carries the number while the expensive reasoning route caches nothing.
- Latency saved and cost saved, not just hits. A high hit rate on requests that were already fast is not a win.
- Cache correctness, which is the metric teams skip. For the semantic layer especially, sample hits and check whether the served answer was actually right for the question asked. Run your golden set through the cache path as well as the cold path and compare scores β if the cached path scores lower, the cache is a quality regression wearing a cost-saving badge. This is an eval, not a dashboard: see Evals.
- Similarity score distribution of hits, with the near-threshold band sampled by a human periodically. Wrong answers cluster just above the line.
- Staleness age of served entries, so you can tell a freshness complaint from a quality complaint.
Never cache: anything time-sensitive or live, anything permission-scoped, anything personalised, refusals and error responses, outputs that failed schema validation, and β across tenant boundaries β anything at all.
Bad to Good to Great
Bad - cache on the question text alone
A dictionary from question string to answer string, process-local, no TTL. It leaks across tenants, survives prompt edits and model changes, serves stale answers after re-indexing, fragments across replicas, and caches errors. It will appear to work perfectly in testing.
Good - a shared exact-match cache with a complete key and a TTL
Redis or Memcached, keyed on the full request identity including tenant, prompt version, model pin, and index version, with a TTL and async selective write-back. This is safe, genuinely effective, and enough for many products. Its limit is that paraphrase always misses, so the hit rate plateaus.
Great - three layers, ordered, with correctness measured
- Prompt ordering fixed so the stable prefix is actually stable and provider-side caching can engage.
- Exact-match layer on a complete versioned key in a shared store, with async selective write-back and request coalescing.
- Semantic layer only where it is defensible β bounded domain, per-tenant namespace, threshold tuned on labelled pairs for precision, time-sensitive and permission-scoped traffic excluded by policy rather than by hope.
- Invalidation by key version for prompt, model, index, and embedding model, plus TTLs for the rest.
- Stampede protection on hot keys.
- Cache correctness in the eval suite, with the cached path scored against the same golden set as the cold path, and near-threshold hits sampled by a human.
The step from Good to Great is not the semantic cache. It is step 6 β being able to prove the cache did not make the product worse.
When to Use
β Cache when:
- Traffic is repetitive, which most real production traffic is far more than test traffic suggests
- A long system prompt, tool catalogue, or shared document repeats on every request, making prefix caching nearly free to adopt
- Cost or time to first token is a binding constraint and the answer for identical input is genuinely stable
- You have a bounded, low-stakes, paraphrase-heavy surface such as an FAQ assistant, for the semantic layer specifically
β Do not cache when:
- The answer depends on the current state of something β balance, status, inventory, availability
- The answer depends on who is asking, either through personalisation or permissions
- You cannot yet enumerate every input that determines the output, because you cannot build a correct key
- You would be adding a semantic layer to a high-stakes domain where a confidently wrong answer costs more than a model call
- The output is a refusal, an error, or failed validation
Common Interview Questions
Q1: What goes into the cache key for an LLM response cache?
Everything that can change the output. Model id including the pinned snapshot, all sampling parameters, the system prompt version, the full message list in order, the tool schema version, the identifiers and version of any retrieval context, and the tenant and user scope. The failure I care most about is omitting tenant or user: two customers asking a byte-identical question have different correct answers, so keying on question text alone serves one customerβs data to another. That is a correctness bug and a security bug at once. I also version the prompt, model, and index into the key so that changing any of them produces a natural miss instead of relying on someone remembering to purge.
Q2: Semantic caching gave us a large hit rate improvement. What would you check before celebrating?
Whether any of those hits were wrong. Similarity is not equivalence β the embedding distance between βcan I cancel within 30 daysβ and βcan I cancel after 30 daysβ is tiny while the correct answers are opposite, so a loose threshold ships confidently wrong answers that arrive fast and fluent. I would tune the threshold against hand-labelled same-answer and different-answer pairs weighted toward negations and qualifier swaps, pick for precision rather than hit rate, sample the hits sitting just above the threshold, and run the golden set through the cache path to see whether the cached path scores below the cold path. I would also confirm the index is namespaced per tenant and that personalised, permission-scoped, and time-sensitive queries are excluded by policy.
Q3: How does prompt ordering affect cost?
Provider-side prefix caching matches from the first token and stops at the first difference, so a single variable token near the top of the prompt invalidates every cacheable token after it. Stable material β system prompt, tool schemas, shared documents, exemplars β goes first; retrieved context and the user turn go last. This is the reverse of what a naive template does, since templates like to interpolate a name, a session id, or todayβs date into the opening line and then append the long shared policy text below. Moving that variable material to the bottom changes no behaviour and makes the whole prefix eligible. It is also why non-deterministic ordering of retrieved chunks quietly destroys hit rate on two layers at once.
Q4: You re-indexed the knowledge base. What happens to your caches?
Every cached answer derived from the old index is potentially wrong, so the index version has to be part of the key β then re-indexing invalidates the affected entries by construction rather than by a purge step someone has to remember mid-release. The same applies to a prompt edit and a model change, which is why release identity and cache keys are the same problem. An embedding model change is the harsher case: old and new vectors are not comparable, so the semantic cache has to be rebuilt rather than reused, and it needs the same migration planning as the retrieval index itself.
Q5: A popular cached answer expires and traffic spikes on the model. What do you do?
That is a cache stampede, and on an LLM path it bites harder than on a database because concurrency on accelerators is the scarce resource, so a hundred simultaneous identical calls degrade latency for everyone, not just the callers. The primary fix is request coalescing β one request computes while the others wait on that single result, enforced with a short-lived per-key lock. Then serve stale while revalidating for entries where mild staleness is better than a latency cliff, and stagger TTLs so entries written together do not expire together. If the upstream is already saturated, shedding or queueing at the edge with a circuit breaker is preferable to piling on retries.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts