Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 19 min read

Prompt and Semantic Caching - Complete Deep Dive

Stage 7 - Production Lesson 32 of 36

Prerequisites: Caching, Embeddings, Tokens and Cost Math Used in: Latency Engineering, GPU Cost Engineering, Deploying AI Features Build it: Lesson 21 - Latency and Cost Engineering implements this as runnable, tested code you can execute offline.


What is Caching in an LLM System?

Caching an LLM call means not paying for generation you have already paid for. The payoff is unusually large here, because unlike a database query a model call is expensive in three currencies at once β€” money, latency, and scarce accelerator capacity β€” so a cache hit returns a saving in all three.

The trap is that β€œLLM caching” names three completely different mechanisms that teams routinely discuss as one thing. They key on different material, they live in different places, and β€” the part that matters β€” they have different correctness properties. One of them cannot return a wrong answer. One of them changes only your bill. One of them can serve a confidently wrong answer to a question nobody has asked before. Treating them as interchangeable is how a cost optimisation becomes an incident.

Real-world analogy: think of a support desk. Exact-match caching is the agent recognising the identical ticket they closed an hour ago and pasting the same reply β€” safe, because it is the same question. Prompt caching is the agent already having the policy manual open on their desk instead of walking to the archive each time β€” same answer, less setup. Semantic caching is the agent deciding your question is β€œbasically the same” as one from last week and reusing that reply. Sometimes that is excellent service. Sometimes the difference between the two questions was the word β€œnot”.


The Three Layers Side by Side

Layer What it keys on Typical hit rate Risk Correctness property
Exact-match response cache A hash of everything that determines the output - model id and version, sampling params, system prompt version, the full message list, tool schemas, retrieval context, tenant and user scope Higher than teams expect on real traffic, because production traffic is far more repetitive than test traffic Low, and entirely about key construction Exact. Identical input gives the response the model already gave. A wrong answer here means your key was missing a component.
Provider-side prompt caching A long stable prefix of the request, matched by the provider from the first token forward Depends almost entirely on prompt ordering, not on query repetition Low. It discounts repeated prefill, it does not change what is generated Neutral. The model still runs. Output is unaffected, only the cost and time to first token change.
Semantic cache An embedding of the query, matched by similarity against stored queries above a threshold Can be high, and that is exactly the seduction High. This is the only layer that can invent a wrong answer Approximate. A hit asserts β€œsimilar enough”, which is not the same claim as β€œequivalent”. The threshold is a correctness knob.

Read the last column top to bottom. That progression β€” exact, neutral, approximate β€” is the whole page.


Layer 1 - Exact-Match Response Caching

Hash the request, store the response, return it on the next identical request. Unglamorous and the layer most teams under-build.

The single engineering decision is what goes into the key, and the rule is absolute: the key must contain every input that could change the output. Anything you leave out is something you are claiming does not matter.

def cache_key(req) -> str:
    return sha256(canonical_json({
        "model": req.model_id,              # family and pinned snapshot, not just "the fast one"
        "params": {                          # sampling changes the output distribution
            "temperature": req.temperature,
            "top_p": req.top_p,
            "max_tokens": req.max_tokens,
            "seed": req.seed,
        },
        "prompt_version": req.prompt_version,    # a prompt edit must miss the cache
        "messages": req.messages,                # the full list, in order, verbatim
        "tools": req.tool_schema_version,        # different tools - different behaviour
        "context_ids": sorted(req.retrieved_chunk_ids),
        "context_version": req.index_version,    # the corpus moved - the answer may move
        "tenant": req.tenant_id,                 # isolation is part of the key
        "locale": req.locale,
    })).hexdigest()

The bug worth naming explicitly. Forget tenant and you will eventually serve one customer’s answer to another customer. That is simultaneously a correctness bug and a data-leak incident, and it is not hypothetical β€” it is the default outcome of keying on β€œthe question text” alone, because two tenants asking β€œwhat is our current discount tier” produce byte-identical questions and different correct answers. The same reasoning applies to user identity on anything personalised and to role on anything permission-scoped.

Two more things belong in the key that people forget because they are not part of the message list: the pinned model snapshot (a provider-side model update changes behaviour with no deploy on your side β€” see Deploying AI Features) and the retrieval index version (the same question over a re-indexed corpus is a different question).

Everything else is ordinary cache engineering, which Caching already covers: where it lives, eviction policy, memory budget. Use a shared store such as Redis or Memcached rather than a per-process dictionary, because per-process caches fragment your hit rate across replicas.


Layer 2 - Provider-Side Prompt Caching

Providers can cache the internal state of a long, stable prefix so that repeat requests sharing that prefix skip most of the prefill work. You do not manage the cache; you make your prompt eligible for it.

The consequence is sharper than it first sounds: prompt order becomes a cost decision. Prefix matching starts at the first token and stops at the first difference, so a single variable token near the top invalidates everything after it.

flowchart TD
    subgraph Stable["Stable prefix - eligible for reuse"]
        P1[System prompt at a pinned version]
        P2[Tool schemas]
        P3[Shared reference document or policy text]
        P4[Few shot exemplars]
    end
    subgraph Variable["Variable tail - prefilled every time"]
        V1[Retrieved chunks for this query]
        V2[Conversation turn from this user]
    end
    P1 --> P2
    P2 --> P3
    P3 --> P4
    P4 --> V1
    V1 --> V2

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class P1,P2,P3,P4 data
    class V1,V2 client

This is the opposite of what a naive template does. The instinctive template interpolates the user’s name, the current timestamp, the session id, or today’s date into the opening line of the system prompt, then appends the long shared policy document underneath. That layout is prefix-hostile: the first line differs on every request, so nothing downstream can be reused. Moving the same variable material to the bottom changes no behaviour and makes the whole document cacheable.

Practical rules:


Layer 3 - Semantic Caching, and Why It Is Dangerous

Embed the incoming query, search stored query embeddings, and if the nearest neighbour is similar enough, return its stored answer without calling the model. Mechanically it is a small retrieval system β€” see Embeddings.

Here is the core lesson of this page: similarity is not equivalence. Cosine similarity measures proximity in a representation trained to place related text near related text. It was never trained to certify that two questions have the same answer. A hit is your system asserting semantic equivalence on the strength of a number that does not mean equivalence.

The threshold is therefore not a tuning parameter, it is a correctness knob. Loosen it and the hit rate climbs while wrong answers start shipping β€” and they ship in the worst possible form, since a cached answer arrives fast, fluent, and previously validated.

The canonical failure is negation and small qualifiers, because the tokens that invert a required answer barely move an embedding:

Stored query Incoming query Embedding distance Correct answer
Is this expense reimbursable Is this expense not reimbursable Tiny Inverted
Can I cancel within 30 days Can I cancel after 30 days Tiny Opposite
What is the fee for domestic transfers What is the fee for international transfers Small Different number
Does the policy cover this Did the 2024 policy cover this Small Possibly different
How do I enable two factor auth How do I disable two factor auth Small Different procedure

Every row is a plausible support question and every row is a production incident if the threshold is loose.

Mitigations, in the order they matter:

  1. Tune the threshold on labelled pairs, and tune it conservatively. Build a set of query pairs hand-labelled as same-answer or different-answer, weighted toward the negation and qualifier cases above. Measure false-hit rate at each candidate threshold and pick for precision, not for hit rate. A semantic cache with a 20 percent hit rate and no wrong answers beats one with a 60 percent hit rate and occasional inversions.
  2. Scope the cache per tenant, and usually per user. A shared semantic cache across tenants compounds the approximation problem with a leak. Namespace the index.
  3. Never semantic-cache personalised or permission-dependent answers. If the correct response depends on who is asking or what they may see, similarity of the question tells you nothing about reusability of the answer.
  4. Exclude anything time-sensitive. Balances, statuses, availability, inventory, β€œlatest” anything. The question is stable and the answer is not.
  5. Prefer narrow domains. Semantic caching works best on a bounded FAQ-shaped surface where paraphrase is common and stakes are low, and worst on open-ended reasoning.
  6. Verify on hit when the stakes justify it. A cheap model or a rule check can confirm the cached answer still addresses the new question β€” see Model Selection and Routing. You give back some of the saving to buy back correctness.
  7. Log every hit with its similarity score and both queries. Without this you cannot audit the layer at all, and you will not know what it got wrong.

The Fall-Through Path

The three layers compose. Cheapest and safest first, model last, write-back on the way out.

flowchart LR
    U[Client request] --> K[Build key from full request identity]
    K --> E[Exact match cache]
    E -->|hit| R[Return stored response]
    E -->|miss| S[Semantic cache lookup per tenant]
    S -->|above threshold| R
    S -->|below threshold or miss| M[Model call using cached stable prefix]
    M --> R
    M --> W[Write back]
    W --> X[Exact cache entry with TTL]
    W --> Y[Semantic index entry with query embedding]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class U client
    class K,R edge
    class E,S,X,Y data
    class M service
    class W async

Two details the diagram encodes deliberately. The write-back is asynchronous β€” a cache write must never sit on the response path, and a failed cache write must never fail the request. And write-back is selective: responses that failed validation, were refused, abstained, or errored do not belong in either cache. Caching a bad answer converts a transient failure into a durable one.


Invalidation, TTLs, and the Stampede

Invalidate by versioning the key, not by remembering to purge. Three things change the correct answer to an unchanged question: the corpus was re-indexed, the prompt was edited, or the model moved. Each one must appear in the key so the change produces a natural miss. The alternative β€” a purge step someone has to remember during a release β€” fails the first time a release is rushed. This is the same argument for release identity made in Deploying AI Features.

Version these into the key explicitly: prompt_version, model_id including the pinned snapshot, index_version, embedding_model_version for the semantic layer, and tool_schema_version. An embedding model change is especially unforgiving, because old stored vectors are not comparable with new ones β€” that cache must be rebuilt, not reused.

TTLs handle the freshness you cannot version. Version keys catch changes you know about; a TTL bounds your exposure to the ones you do not. Set it from how stale an answer may acceptably be: long for stable reference material, short for anything operational, zero for anything live.

The stampede. A popular key expires or goes cold and every concurrent request for it misses at once, all of them calling the model with the same input. On a GPU-bound path this is worse than a database stampede, because concurrency is the scarce resource β€” see GPU Cost Engineering. Standard mitigations transfer: a per-key lock so one request computes while the others wait on the result, serve-stale-while-revalidating for entries where slight staleness beats a latency spike, and staggered TTLs so entries written together do not expire together. Coalescing concurrent identical requests into one model call is the single highest-value version of this.


Measuring the Cache

Hit rate is the metric everyone reports and it is only half the instrumentation. Track per layer:

Never cache: anything time-sensitive or live, anything permission-scoped, anything personalised, refusals and error responses, outputs that failed schema validation, and β€” across tenant boundaries β€” anything at all.


Bad to Good to Great

Bad - cache on the question text alone

A dictionary from question string to answer string, process-local, no TTL. It leaks across tenants, survives prompt edits and model changes, serves stale answers after re-indexing, fragments across replicas, and caches errors. It will appear to work perfectly in testing.

Good - a shared exact-match cache with a complete key and a TTL

Redis or Memcached, keyed on the full request identity including tenant, prompt version, model pin, and index version, with a TTL and async selective write-back. This is safe, genuinely effective, and enough for many products. Its limit is that paraphrase always misses, so the hit rate plateaus.

Great - three layers, ordered, with correctness measured

  1. Prompt ordering fixed so the stable prefix is actually stable and provider-side caching can engage.
  2. Exact-match layer on a complete versioned key in a shared store, with async selective write-back and request coalescing.
  3. Semantic layer only where it is defensible β€” bounded domain, per-tenant namespace, threshold tuned on labelled pairs for precision, time-sensitive and permission-scoped traffic excluded by policy rather than by hope.
  4. Invalidation by key version for prompt, model, index, and embedding model, plus TTLs for the rest.
  5. Stampede protection on hot keys.
  6. Cache correctness in the eval suite, with the cached path scored against the same golden set as the cold path, and near-threshold hits sampled by a human.

The step from Good to Great is not the semantic cache. It is step 6 β€” being able to prove the cache did not make the product worse.


When to Use

βœ… Cache when:

❌ Do not cache when:


Common Interview Questions

Q1: What goes into the cache key for an LLM response cache?

Everything that can change the output. Model id including the pinned snapshot, all sampling parameters, the system prompt version, the full message list in order, the tool schema version, the identifiers and version of any retrieval context, and the tenant and user scope. The failure I care most about is omitting tenant or user: two customers asking a byte-identical question have different correct answers, so keying on question text alone serves one customer’s data to another. That is a correctness bug and a security bug at once. I also version the prompt, model, and index into the key so that changing any of them produces a natural miss instead of relying on someone remembering to purge.

Q2: Semantic caching gave us a large hit rate improvement. What would you check before celebrating?

Whether any of those hits were wrong. Similarity is not equivalence β€” the embedding distance between β€œcan I cancel within 30 days” and β€œcan I cancel after 30 days” is tiny while the correct answers are opposite, so a loose threshold ships confidently wrong answers that arrive fast and fluent. I would tune the threshold against hand-labelled same-answer and different-answer pairs weighted toward negations and qualifier swaps, pick for precision rather than hit rate, sample the hits sitting just above the threshold, and run the golden set through the cache path to see whether the cached path scores below the cold path. I would also confirm the index is namespaced per tenant and that personalised, permission-scoped, and time-sensitive queries are excluded by policy.

Q3: How does prompt ordering affect cost?

Provider-side prefix caching matches from the first token and stops at the first difference, so a single variable token near the top of the prompt invalidates every cacheable token after it. Stable material β€” system prompt, tool schemas, shared documents, exemplars β€” goes first; retrieved context and the user turn go last. This is the reverse of what a naive template does, since templates like to interpolate a name, a session id, or today’s date into the opening line and then append the long shared policy text below. Moving that variable material to the bottom changes no behaviour and makes the whole prefix eligible. It is also why non-deterministic ordering of retrieved chunks quietly destroys hit rate on two layers at once.

Q4: You re-indexed the knowledge base. What happens to your caches?

Every cached answer derived from the old index is potentially wrong, so the index version has to be part of the key β€” then re-indexing invalidates the affected entries by construction rather than by a purge step someone has to remember mid-release. The same applies to a prompt edit and a model change, which is why release identity and cache keys are the same problem. An embedding model change is the harsher case: old and new vectors are not comparable, so the semantic cache has to be rebuilt rather than reused, and it needs the same migration planning as the retrieval index itself.

Q5: A popular cached answer expires and traffic spikes on the model. What do you do?

That is a cache stampede, and on an LLM path it bites harder than on a database because concurrency on accelerators is the scarce resource, so a hundred simultaneous identical calls degrade latency for everyone, not just the callers. The primary fix is request coalescing β€” one request computes while the others wait on that single result, enforced with a short-lived per-key lock. Then serve stale while revalidating for entries where mild staleness is better than a latency cliff, and stagger TTLs so entries written together do not expire together. If the upstream is already saturated, shedding or queueing at the edge with a circuit breaker is preferable to piling on retries.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access