Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 22 min read

Latency Engineering for LLM Apps - Complete Deep Dive

Stage 7 - Production Lesson 31 of 36

Prerequisites: Inference Serving, Performance Metrics, Model Selection and Routing Used in: Prompt and Semantic Caching, Tracing and Observability for LLM Apps, ChatGPT Build it: Lesson 21 - Latency and Cost Engineering implements this as runnable, tested code you can execute offline.


What is Latency Engineering for an LLM App?

Latency engineering is the practice of deciding which clock your users are actually reading, budgeting that clock across every stage of the request, and then spending effort on the stage that dominates rather than the stage you find interesting.

The trap is specific and almost universal: teams track one number called β€œlatency”, see it move, and start optimising the model call. But an LLM request has at least two clocks that fail for different reasons and respond to different fixes:

For streamed output these are not two views of one problem. Perceived latency is dominated by TTFT, because once tokens are arriving faster than the user reads, extra completion time is nearly invisible. For non-streamed output, or for anything a downstream program consumes, total time is the only number that matters. Budget them separately or you will optimise the wrong half.

Real-world analogy: Two ways to watch a film. One downloads completely before it plays; the other starts playing in two seconds and buffers as it goes. Both deliver the same bytes in roughly the same total time, and only one feels broken. Streaming does not make the system faster β€” it changes which clock the user is reading. Most of this page is that observation applied methodically.


One Number Is the Wrong Metric

Metric What it measures What the user feels Which phase owns it
Time to first token Request start until the first visible token Whether the system is alive Network, auth, retrieval, prompt assembly, prefill
Inter-token latency Gap between successive streamed tokens Whether reading feels natural or stuttery Decode, plus fleet load and batch size
Total completion time Request start until the final token How long the task took Output length, almost entirely
Time to first useful token First token that carries information Real responsiveness Prompt design - a long preamble wastes the win
End-to-end task time Multi-step or agent run from start to finish Whether the feature is usable at all Number of steps and each step’s tail

Two deserve emphasis. Time to first useful token is the honest version of TTFT: if your format makes the model emit 40 tokens restating the question before the answer begins, you bought a fast TTFT and shipped a slow experience. And end-to-end task time is the one agent products get wrong, because per-step latency looks fine while the run takes a minute. Track all of them as distributions, per route, with the model and prompt version attached. Percentile discipline generally is in Performance Metrics; getting per-stage timings out of a real request needs the span instrumentation in Tracing and Observability for LLM Apps.


The Latency Budget for a Real RAG Request

Before optimising anything, decompose. A RAG request is not β€œone model call” β€” it is a chain, and most of the chain is yours, not the provider’s.

flowchart LR
    C[Client request] --> NET[Network and TLS]
    NET --> AUTH[Auth and rate limit check]
    AUTH --> EMB[Embed the query]
    EMB --> VS[Vector search]
    VS --> RR[Rerank the candidates]
    RR --> ASM[Prompt assembly]
    ASM --> PRE[Model prefill - owns TTFT]
    PRE --> DEC[Model decode - owns total time]
    DEC --> VAL[Output validation]
    VAL --> OUT[Response complete]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class C client
    class NET,AUTH edge
    class EMB,VS,RR service
    class ASM,VAL async
    class PRE,DEC data
    class OUT edge

Every number in the table below is an ASSUMPTION invented to make the arithmetic readable. None of it is a benchmark, a quote, or a measurement of any real model, index, or reranker. The point is the shape of the decomposition and the method for finding your own dominant stage. Measure your own stages before believing any of it.

Stage ASSUMED contribution Counts toward
Network and TLS 30 ms TTFT and total
Auth and rate limit check 5 ms TTFT and total
Embed the query 40 ms TTFT and total
Vector search 25 ms TTFT and total
Rerank the candidates 180 ms TTFT and total
Prompt assembly 5 ms TTFT and total
Model prefill 350 ms TTFT and total
Subtotal - TTFT 635 ms what the user waits for
Model decode - ASSUMED 300 output tokens at an ASSUMED 40 tokens per second 7,500 ms total only
Output validation 20 ms total only
Total completion time about 8.2 s Β 

Two readings, and both are the reason to do this exercise.

The reranker is 28 percent of TTFT in this assumed budget, and it is not the model.
πŸ’‘ A reranker is a second-pass model that re-scores the handful of passages search returned, so the best ones end up at the top.
A team that spent a week on prompt compression to shave prefill was working on the wrong stage. The better question is whether the reranker needs to see 50 candidates or 15, or whether it needs to run at all on a cache hit. The same applies to a cold embedding cache, a cross-region hop to the vector store, or an auth call that does an uncached database lookup. Teams reflexively optimise the model call while a reranker or a cold cache is the actual cost β€” see Hybrid Search and Reranking for what the reranker buys you, so the trade is deliberate.

Decode dominates total time β€” the overwhelming majority of it in this assumed budget β€” and nothing in the retrieval stack can touch it. Which brings us to the single largest lever.


Output Length Is the Dominant Driver

Prefill processes the whole prompt in one parallel pass. Decode produces exactly one token per forward pass, strictly sequentially, because token N depends on token N-1. The mechanism is in Inference Serving; the consequence for your product is blunt:

Total completion time is roughly linear in output length, and you control output length.

That makes two things you might have filed under β€œprompt style” into genuine latency controls:

The corollary is uncomfortable for reasoning-heavy designs: chain-of-thought, self-critique passes, and verbose scratchpads buy quality with sequential tokens. That may well be the right trade β€” but price it in milliseconds, keep reasoning out of the user-visible stream where the format allows, and do not let it creep onto a latency-critical path unnoticed.


Streaming Is the Highest-Leverage Perceived-Latency Win

If you ship one thing from this page, ship streaming. It does not reduce total time by a millisecond and it routinely makes an interface feel several times faster, because it converts an 8-second blank screen into a 0.6-second wait followed by text the user reads while it arrives.


The Techniques, Roughly Ordered by Leverage

1. Cache at every layer. The only optimisation that can take a stage to zero. Exact-match response caching, provider-side prefix caching, and a semantic cache with its correctness caveats are all in Prompt and Semantic Caching. Also cache the stages nobody thinks of: query embeddings, reranker scores for a repeated query, and auth or entitlement lookups.

2. Cut retrieved context. Fewer and shorter chunks means less prefill and a lower TTFT, plus a cheaper request. Measure recall as you cut β€” you are trading retrieval quality for time, and the only way to know the price is an eval run.

3. Put a smaller or faster model on latency-critical paths. Classification, routing, query rewriting, and moderation rarely need your largest model. Routing by difficulty is the same machinery as cost routing in Model Selection and Routing, with latency as the objective instead of price.

4. Parallelise independent steps. Most chains are written sequentially out of habit, not necessity. A moderation check does not depend on retrieval; two retrievers over different corpora do not depend on each other; a metadata lookup does not depend on the embedding. Fan those out and join before the model call: the serial path costs the sum of its stages, the parallel path costs the maximum. The price is wasted work when the moderation check rejects after retrieval already ran, which is usually trivial next to the time saved.

5. Speculative decoding. Decode is sequential and memory-bandwidth-bound, so the attack is to get more than one token out of each expensive pass. A small draft model proposes several tokens ahead, and the large model verifies all of them in a single forward pass, accepting the longest run that matches what it would have produced itself and discarding the rest. Because rejected tokens are thrown away, the output distribution is unchanged β€” this is a latency optimisation, not a quality trade. Several tokens can be accepted per verification pass, so total completion time falls. Two caveats: it does nothing for TTFT, since prefill is untouched; and the gain depends entirely on how predictable your text is, so a self-hosted deployment must measure it on real traffic rather than assume it.
πŸ’‘ A cheap model guesses the next few words, the expensive model checks the whole guess in one go, and anything it would not have written itself gets thrown away.

6. Prompt caching for a stable prefix. Order your prompt so the unchanging material β€” system prompt, tool schemas, few-shot examples β€” comes first and the volatile material comes last. This cuts prefill and therefore TTFT directly. A per-request timestamp or session ID at the top of the prompt silently destroys the match for every request.
πŸ’‘ The provider only gets to skip re-reading the start of your prompt while that part is character-for-character what it saw last time.

7. Prefetch and precompute where the request is predictable. If a user opening a document almost always asks for a summary next, start generating it on open. If a support agent opens a ticket, warm the retrieval for that customer. Prefetching trades wasted compute on wrong guesses for a near-zero perceived wait on right ones, so it is worth it exactly when your hit rate is high and the wasted call is cheap.


Technique Scorecard

Technique Affects Cost
Exact-match response cache Both - removes the call entirely Cache infrastructure and strict key discipline; staleness risk
Provider-side prefix caching TTFT Constrains prompt ordering; nothing else
Semantic cache Both Real correctness risk from near-miss hits
Cut retrieved context TTFT mainly Retrieval recall; must be eval-gated
Shrink or skip the reranker TTFT Ranking quality on ambiguous queries
Smaller or faster model Both Quality; needs an eval gate per route
Parallelise independent steps TTFT Concurrency complexity and some wasted work
Speculative decoding Total completion time Draft model memory and compute; workload-dependent gain
Max-token cap and a brevity instruction Total completion time Truncation risk; must be monitored
Streaming Perceived latency only Client and transport work; harder error handling
Prefetch and precompute Both, when the guess is right Wasted compute on wrong guesses
Deadline-aware fallback The tail specifically Lower quality on the fallback path
Remove network hops and colocate TTFT Deployment topology constraints

Tail Latency Is What Users Feel

A median is a comfortable number that describes nobody in particular. p95 and p99 are the experience, because a user issuing twenty requests in a session will hit the tail, and the requests that stall are disproportionately the important ones β€” long documents, complex questions, the demo.
πŸ’‘ p95 means 95 out of every 100 requests finished faster than this number, so p50 is the typical request and p99 is the worst one in a hundred.

For LLM apps the tail has causes worth naming individually: a cold cache, a retry after a provider timeout, an unusually long generation. The rest sit on the serving side - fleet queueing under load, a cold replica still warming, and a slow external tool inside an agent loop. Each is fixed differently, so alert on the percentile and then attribute with traces.

Then the multi-step problem, which is the reason agent products feel slow even when every dashboard looks green. Each step’s tail multiplies the chance the whole run feels slow.

The numbers here are ASSUMPTIONS chosen to show the arithmetic, not measurements. Assume a 6-step agent run where each step independently has a 5 percent chance of exceeding its budget. The chance that at least one step stalls is 1 - 0.95^6, which is about 26 percent. A per-step p95 that looks respectable in isolation has become a roughly one-in-four chance that any given run is slow.

Three consequences. Fewer steps is a latency feature, so collapse steps where a single call can do the work. Cap the steps, because an unbounded loop has an unbounded tail. And show progress for anything genuinely long, since a 40-second run with visible intermediate state is tolerable and a 40-second spinner is not.


Deadlines, Timeouts, and Fallbacks

A timeout per call is table stakes and not enough, because independent per-call timeouts sum: five stages capped at five seconds each is a twenty-five-second worst case nobody signed up for. What you want is one deadline for the request, propagated down the chain, with every stage asking β€œhow much budget is left” before it starts.

flowchart TD
    R[Request arrives with a total deadline] --> A[Allocate a budget to each stage]
    A --> CA[Cache lookup]
    CA --> HIT[Cache hit - return the stored answer]
    CA --> RE[Retrieval with its own stage timeout]
    RE --> CK[Deadline check before the model call]
    CK --> FULL[Ample budget - call the primary model]
    CK --> FAST[Budget at risk - call the smaller faster model]
    CK --> DEG[No budget left - return a cached or degraded answer]
    FULL --> S[Stream tokens to the client]
    FAST --> S

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class R client
    class A,CK async
    class CA,RE service
    class HIT,DEG data
    class FULL,FAST service
    class S edge
def handle(req, deadline_ms: int):
    clock = Deadline(deadline_ms)               # one budget for the whole request
    if hit := cache.get(req):
        return hit
    ctx = retrieve(req, timeout_ms=clock.slice(0.25))
    if clock.remaining_ms() < MODEL_MIN_BUDGET_MS:
        return degraded_answer(req, ctx)        # explicit, logged, measurable
    model = PRIMARY if clock.remaining_ms() > PRIMARY_BUDGET_MS else FAST
    return generate(model, req, ctx, timeout_ms=clock.remaining_ms())

Four rules that make this hold up in production. Never let the sum of stage timeouts exceed the request deadline. Only retry when the remaining budget can actually absorb the retry, otherwise you are converting a slow request into a slower failure β€” the backoff mechanics are in Retries and Backoff. Stop calling a provider that is already failing, using the pattern in Circuit Breaker, because queueing behind a dead dependency spends the whole budget on nothing. And count the fallbacks β€” a fallback path that fires on 30 percent of traffic is not a safety valve, it is your production system, and it should be evaluated like one.


Perceived Latency - The UX Levers

These are not consolation prizes for failing to optimise. They routinely buy more apparent speed than a week of backend work.


Bad to Good to Great

Bad - one latency dashboard and a faster model

A single average latency chart. When it rises, swap in a faster model and hope. You cannot tell whether TTFT or completion time moved, you have no per-stage attribution, so the reranker or the cold cache that is actually responsible stays invisible, and the model swap silently costs quality you never measured.

Good - streaming, a max-token cap, and a response cache

Real wins, in the right order, and enough to make most features feel acceptable. What is still missing: no per-stage budget, so the next regression is again a guessing game. Per-call timeouts with no request deadline, so the worst case is the sum of every stage. Nothing specific about the tail, and no defined behaviour when a dependency is slow rather than down.

Great - budgeted stages, deadline-aware execution, and percentile SLOs per route

  1. TTFT and total completion time budgeted and measured separately, per route, with traces giving per-stage timings.
  2. Percentile SLOs, not averages, with p95 and p99 alerting and attribution by stage.
  3. Caching at every layer, including the embedding, reranker, and auth lookups teams forget.
  4. Output length controlled deliberately by route caps and format instructions, with independent stages parallelised and unnecessary hops removed.
  5. One deadline propagated through the chain, with a faster model and a degraded answer as defined fallbacks rather than accidents, and fallback and truncation rates tracked as product metrics so silent degradation cannot hide.
  6. Perceived-latency work treated as first-class β€” streaming, optimistic rendering, progressive disclosure, real progress for long runs.

When to Use

βœ… Invest in latency engineering when:

❌ Do not chase latency when:


Common Interview Questions

Q1: Your LLM feature feels slow. Where do you start?

By splitting the number, then decomposing the request. First, is TTFT the problem or is total completion time the problem, because they have different causes and different fixes β€” and if output is streamed, TTFT is what the user perceives. Then I get per-stage timings from traces for a real request: network, auth, query embedding, vector search, reranking, prompt assembly, prefill, decode, validation. The dominant stage is frequently not the model call β€” a reranker scoring too many candidates, a cold embedding cache, or a cross-region hop to the vector store are common. Only once I know the dominant stage do I pick a technique, because optimising anything else is measurable work with unmeasurable benefit.

Q2: What is the difference between TTFT and total completion time, and which should you optimise?

TTFT is everything up to and including prefill, which processes the whole prompt in one parallel pass. Total completion time is dominated by decode, which emits one token per forward pass sequentially, so it scales with output length. For streamed output to a human, optimise TTFT first β€” once tokens arrive faster than the user reads, extra completion time is close to invisible. For output a program consumes, or any non-streamed response, total time is the only clock that matters. The biggest lever there is output length, so a max-token cap and an instruction to be concise are genuine latency controls rather than style preferences.

Q3: Explain speculative decoding.

It attacks the fact that decode produces one token per expensive pass over the weights. A small draft model proposes several tokens ahead, cheaply. The large model then verifies that whole proposed span in a single forward pass and accepts the longest prefix matching what it would have generated itself, discarding the rest. Because rejected tokens are thrown away, the output distribution is identical to running the large model alone β€” it is purely a latency win, not a quality trade. Several tokens per verification pass means lower total completion time. It does nothing for TTFT since prefill is unchanged, and the acceptance rate depends on how predictable the text is, so the gain has to be measured on your own traffic.

Q4: Why do p95 and p99 matter more than the average here, and what does that mean for agents?

Because users experience individual requests, not a distribution, and a session of twenty requests will land in the tail. The tail also has distinct causes worth attributing separately: cold caches, retries, unusually long generations, fleet queueing, cold replicas, and slow external tools. In a multi-step agent it compounds β€” each step’s tail is another chance for the run to feel slow, so with an assumed six steps each having a 5 percent chance of exceeding budget, roughly a quarter of runs stall somewhere. The responses are to reduce step count, cap the loop, and show real intermediate progress so a long run stays tolerable.

Q5: How do you stop a slow dependency from blowing your latency SLO?

One deadline for the request, propagated down the chain, instead of independent per-call timeouts that sum to a worst case nobody agreed to. Every stage gets a budget and checks the remaining budget before starting. When the budget is at risk before the model call, I degrade deliberately: route to a smaller faster model, or return a cached or reduced answer, rather than letting the request run long and then fail. Retries only happen when the remaining budget can absorb one, and a circuit breaker stops me queueing behind a provider that is already failing. Critically, I count how often the fallback fires β€” if it is firing on a large share of traffic it is not a safety valve, it is the production path, and it needs an eval of its own.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access