Latency Engineering for LLM Apps - Complete Deep Dive
Prerequisites: Inference Serving, Performance Metrics, Model Selection and Routing Used in: Prompt and Semantic Caching, Tracing and Observability for LLM Apps, ChatGPT Build it: Lesson 21 - Latency and Cost Engineering implements this as runnable, tested code you can execute offline.
What is Latency Engineering for an LLM App?
Latency engineering is the practice of deciding which clock your users are actually reading, budgeting that clock across every stage of the request, and then spending effort on the stage that dominates rather than the stage you find interesting.
The trap is specific and almost universal: teams track one number called βlatencyβ, see it move, and start optimising the model call. But an LLM request has at least two clocks that fail for different reasons and respond to different fixes:
- Time to first token (TTFT) β how long the user stares at nothing. Set by everything before and including prefill.
π‘ Prefill is the single pass where the model reads your whole prompt before writing anything, and decode is the writing that follows, one word at a time. - Total completion time β how long until the whole answer exists. Dominated by decode, which is sequential.
For streamed output these are not two views of one problem. Perceived latency is dominated by TTFT, because once tokens are arriving faster than the user reads, extra completion time is nearly invisible. For non-streamed output, or for anything a downstream program consumes, total time is the only number that matters. Budget them separately or you will optimise the wrong half.
Real-world analogy: Two ways to watch a film. One downloads completely before it plays; the other starts playing in two seconds and buffers as it goes. Both deliver the same bytes in roughly the same total time, and only one feels broken. Streaming does not make the system faster β it changes which clock the user is reading. Most of this page is that observation applied methodically.
One Number Is the Wrong Metric
| Metric | What it measures | What the user feels | Which phase owns it |
|---|---|---|---|
| Time to first token | Request start until the first visible token | Whether the system is alive | Network, auth, retrieval, prompt assembly, prefill |
| Inter-token latency | Gap between successive streamed tokens | Whether reading feels natural or stuttery | Decode, plus fleet load and batch size |
| Total completion time | Request start until the final token | How long the task took | Output length, almost entirely |
| Time to first useful token | First token that carries information | Real responsiveness | Prompt design - a long preamble wastes the win |
| End-to-end task time | Multi-step or agent run from start to finish | Whether the feature is usable at all | Number of steps and each stepβs tail |
Two deserve emphasis. Time to first useful token is the honest version of TTFT: if your format makes the model emit 40 tokens restating the question before the answer begins, you bought a fast TTFT and shipped a slow experience. And end-to-end task time is the one agent products get wrong, because per-step latency looks fine while the run takes a minute. Track all of them as distributions, per route, with the model and prompt version attached. Percentile discipline generally is in Performance Metrics; getting per-stage timings out of a real request needs the span instrumentation in Tracing and Observability for LLM Apps.
The Latency Budget for a Real RAG Request
Before optimising anything, decompose. A RAG request is not βone model callβ β it is a chain, and most of the chain is yours, not the providerβs.
flowchart LR
C[Client request] --> NET[Network and TLS]
NET --> AUTH[Auth and rate limit check]
AUTH --> EMB[Embed the query]
EMB --> VS[Vector search]
VS --> RR[Rerank the candidates]
RR --> ASM[Prompt assembly]
ASM --> PRE[Model prefill - owns TTFT]
PRE --> DEC[Model decode - owns total time]
DEC --> VAL[Output validation]
VAL --> OUT[Response complete]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class C client
class NET,AUTH edge
class EMB,VS,RR service
class ASM,VAL async
class PRE,DEC data
class OUT edge
Every number in the table below is an ASSUMPTION invented to make the arithmetic readable. None of it is a benchmark, a quote, or a measurement of any real model, index, or reranker. The point is the shape of the decomposition and the method for finding your own dominant stage. Measure your own stages before believing any of it.
| Stage | ASSUMED contribution | Counts toward |
|---|---|---|
| Network and TLS | 30 ms | TTFT and total |
| Auth and rate limit check | 5 ms | TTFT and total |
| Embed the query | 40 ms | TTFT and total |
| Vector search | 25 ms | TTFT and total |
| Rerank the candidates | 180 ms | TTFT and total |
| Prompt assembly | 5 ms | TTFT and total |
| Model prefill | 350 ms | TTFT and total |
| Subtotal - TTFT | 635 ms | what the user waits for |
| Model decode - ASSUMED 300 output tokens at an ASSUMED 40 tokens per second | 7,500 ms | total only |
| Output validation | 20 ms | total only |
| Total completion time | about 8.2 s | Β |
Two readings, and both are the reason to do this exercise.
The reranker is 28 percent of TTFT in this assumed budget, and it is not the model.
π‘ A reranker is a second-pass model that re-scores the handful of passages search returned, so the best ones end up at the top.
A team that spent a week on prompt compression to shave prefill was working on the wrong stage. The better question is whether the reranker needs to see 50 candidates or 15, or whether it needs to run at all on a cache hit. The same applies to a cold embedding cache, a cross-region hop to the vector store, or an auth call that does an uncached database lookup. Teams reflexively optimise the model call while a reranker or a cold cache is the actual cost β see Hybrid Search and Reranking for what the reranker buys you, so the trade is deliberate.
Decode dominates total time β the overwhelming majority of it in this assumed budget β and nothing in the retrieval stack can touch it. Which brings us to the single largest lever.
Output Length Is the Dominant Driver
Prefill processes the whole prompt in one parallel pass. Decode produces exactly one token per forward pass, strictly sequentially, because token N depends on token N-1. The mechanism is in Inference Serving; the consequence for your product is blunt:
Total completion time is roughly linear in output length, and you control output length.
That makes two things you might have filed under βprompt styleβ into genuine latency controls:
- Asking for brevity works. βAnswer in at most three sentencesβ, βreturn only the JSON objectβ, βno preamble, no summary of the questionβ β each removes tokens that each cost a sequential forward pass. This is the rare optimisation that cuts latency and cost together, with no infrastructure work.
- A max-token cap is a latency guardrail, not a safety net. Without one, a single pathological generation can run until the context ends. Cap per route, sized from the longest legitimate answer you have observed, and treat a truncation as a monitored event rather than a silent one.
The corollary is uncomfortable for reasoning-heavy designs: chain-of-thought, self-critique passes, and verbose scratchpads buy quality with sequential tokens. That may well be the right trade β but price it in milliseconds, keep reasoning out of the user-visible stream where the format allows, and do not let it creep onto a latency-critical path unnoticed.
Streaming Is the Highest-Leverage Perceived-Latency Win
If you ship one thing from this page, ship streaming. It does not reduce total time by a millisecond and it routinely makes an interface feel several times faster, because it converts an 8-second blank screen into a 0.6-second wait followed by text the user reads while it arrives.
- Transport is server-sent events for the common one-directional case, or WebSockets when you need bidirectional traffic such as interruption and mid-stream tool confirmation β see WebSockets.
- Streaming changes your error handling. A failure now arrives after the user has read half an answer, so you need a way to retract or amend a partial response rather than a clean 500.
- Streaming and strict output validation pull in opposite directions. You cannot validate a JSON object you have not finished receiving. Practical split: stream prose to humans, buffer-and-validate anything a program will parse, and see Structured Outputs for the partial-parse options.
- If a downstream program is the consumer, streaming buys nothing. Do not add the complexity for a machine that waits for the last token anyway. The ChatGPT design covers the fleet-level mechanics of streaming to millions of connections; this page is about where streaming sits in your budget.
The Techniques, Roughly Ordered by Leverage
1. Cache at every layer. The only optimisation that can take a stage to zero. Exact-match response caching, provider-side prefix caching, and a semantic cache with its correctness caveats are all in Prompt and Semantic Caching. Also cache the stages nobody thinks of: query embeddings, reranker scores for a repeated query, and auth or entitlement lookups.
2. Cut retrieved context. Fewer and shorter chunks means less prefill and a lower TTFT, plus a cheaper request. Measure recall as you cut β you are trading retrieval quality for time, and the only way to know the price is an eval run.
3. Put a smaller or faster model on latency-critical paths. Classification, routing, query rewriting, and moderation rarely need your largest model. Routing by difficulty is the same machinery as cost routing in Model Selection and Routing, with latency as the objective instead of price.
4. Parallelise independent steps. Most chains are written sequentially out of habit, not necessity. A moderation check does not depend on retrieval; two retrievers over different corpora do not depend on each other; a metadata lookup does not depend on the embedding. Fan those out and join before the model call: the serial path costs the sum of its stages, the parallel path costs the maximum. The price is wasted work when the moderation check rejects after retrieval already ran, which is usually trivial next to the time saved.
5. Speculative decoding. Decode is sequential and memory-bandwidth-bound, so the attack is to get more than one token out of each expensive pass. A small draft model proposes several tokens ahead, and the large model verifies all of them in a single forward pass, accepting the longest run that matches what it would have produced itself and discarding the rest. Because rejected tokens are thrown away, the output distribution is unchanged β this is a latency optimisation, not a quality trade. Several tokens can be accepted per verification pass, so total completion time falls. Two caveats: it does nothing for TTFT, since prefill is untouched; and the gain depends entirely on how predictable your text is, so a self-hosted deployment must measure it on real traffic rather than assume it.
π‘ A cheap model guesses the next few words, the expensive model checks the whole guess in one go, and anything it would not have written itself gets thrown away.
6. Prompt caching for a stable prefix. Order your prompt so the unchanging material β system prompt, tool schemas, few-shot examples β comes first and the volatile material comes last. This cuts prefill and therefore TTFT directly. A per-request timestamp or session ID at the top of the prompt silently destroys the match for every request.
π‘ The provider only gets to skip re-reading the start of your prompt while that part is character-for-character what it saw last time.
7. Prefetch and precompute where the request is predictable. If a user opening a document almost always asks for a summary next, start generating it on open. If a support agent opens a ticket, warm the retrieval for that customer. Prefetching trades wasted compute on wrong guesses for a near-zero perceived wait on right ones, so it is worth it exactly when your hit rate is high and the wasted call is cheap.
Technique Scorecard
| Technique | Affects | Cost |
|---|---|---|
| Exact-match response cache | Both - removes the call entirely | Cache infrastructure and strict key discipline; staleness risk |
| Provider-side prefix caching | TTFT | Constrains prompt ordering; nothing else |
| Semantic cache | Both | Real correctness risk from near-miss hits |
| Cut retrieved context | TTFT mainly | Retrieval recall; must be eval-gated |
| Shrink or skip the reranker | TTFT | Ranking quality on ambiguous queries |
| Smaller or faster model | Both | Quality; needs an eval gate per route |
| Parallelise independent steps | TTFT | Concurrency complexity and some wasted work |
| Speculative decoding | Total completion time | Draft model memory and compute; workload-dependent gain |
| Max-token cap and a brevity instruction | Total completion time | Truncation risk; must be monitored |
| Streaming | Perceived latency only | Client and transport work; harder error handling |
| Prefetch and precompute | Both, when the guess is right | Wasted compute on wrong guesses |
| Deadline-aware fallback | The tail specifically | Lower quality on the fallback path |
| Remove network hops and colocate | TTFT | Deployment topology constraints |
Tail Latency Is What Users Feel
A median is a comfortable number that describes nobody in particular. p95 and p99 are the experience, because a user issuing twenty requests in a session will hit the tail, and the requests that stall are disproportionately the important ones β long documents, complex questions, the demo.
π‘ p95 means 95 out of every 100 requests finished faster than this number, so p50 is the typical request and p99 is the worst one in a hundred.
For LLM apps the tail has causes worth naming individually: a cold cache, a retry after a provider timeout, an unusually long generation. The rest sit on the serving side - fleet queueing under load, a cold replica still warming, and a slow external tool inside an agent loop. Each is fixed differently, so alert on the percentile and then attribute with traces.
Then the multi-step problem, which is the reason agent products feel slow even when every dashboard looks green. Each stepβs tail multiplies the chance the whole run feels slow.
The numbers here are ASSUMPTIONS chosen to show the arithmetic, not measurements. Assume a 6-step agent run where each step independently has a 5 percent chance of exceeding its budget. The chance that at least one step stalls is
1 - 0.95^6, which is about 26 percent. A per-step p95 that looks respectable in isolation has become a roughly one-in-four chance that any given run is slow.
Three consequences. Fewer steps is a latency feature, so collapse steps where a single call can do the work. Cap the steps, because an unbounded loop has an unbounded tail. And show progress for anything genuinely long, since a 40-second run with visible intermediate state is tolerable and a 40-second spinner is not.
Deadlines, Timeouts, and Fallbacks
A timeout per call is table stakes and not enough, because independent per-call timeouts sum: five stages capped at five seconds each is a twenty-five-second worst case nobody signed up for. What you want is one deadline for the request, propagated down the chain, with every stage asking βhow much budget is leftβ before it starts.
flowchart TD
R[Request arrives with a total deadline] --> A[Allocate a budget to each stage]
A --> CA[Cache lookup]
CA --> HIT[Cache hit - return the stored answer]
CA --> RE[Retrieval with its own stage timeout]
RE --> CK[Deadline check before the model call]
CK --> FULL[Ample budget - call the primary model]
CK --> FAST[Budget at risk - call the smaller faster model]
CK --> DEG[No budget left - return a cached or degraded answer]
FULL --> S[Stream tokens to the client]
FAST --> S
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class R client
class A,CK async
class CA,RE service
class HIT,DEG data
class FULL,FAST service
class S edge
def handle(req, deadline_ms: int):
clock = Deadline(deadline_ms) # one budget for the whole request
if hit := cache.get(req):
return hit
ctx = retrieve(req, timeout_ms=clock.slice(0.25))
if clock.remaining_ms() < MODEL_MIN_BUDGET_MS:
return degraded_answer(req, ctx) # explicit, logged, measurable
model = PRIMARY if clock.remaining_ms() > PRIMARY_BUDGET_MS else FAST
return generate(model, req, ctx, timeout_ms=clock.remaining_ms())
Four rules that make this hold up in production. Never let the sum of stage timeouts exceed the request deadline. Only retry when the remaining budget can actually absorb the retry, otherwise you are converting a slow request into a slower failure β the backoff mechanics are in Retries and Backoff. Stop calling a provider that is already failing, using the pattern in Circuit Breaker, because queueing behind a dead dependency spends the whole budget on nothing. And count the fallbacks β a fallback path that fires on 30 percent of traffic is not a safety valve, it is your production system, and it should be evaluated like one.
Perceived Latency - The UX Levers
These are not consolation prizes for failing to optimise. They routinely buy more apparent speed than a week of backend work.
- Streaming, as above. The single biggest one.
- Optimistic UI. Render the userβs message, the skeleton of the answer, and the citations frame immediately. The wait now happens inside a live interface rather than a frozen one.
- Progressive disclosure. Return the direct answer first and stream supporting detail after. A user who got what they needed in the first sentence has stopped waiting, whatever the total time says.
- Intermediate progress for long agent runs. βSearching the policy corpusβ, βchecking three sourcesβ, βdraftingβ β real step labels from real spans, not a fake progress bar. This is also why step-level traces are worth emitting for their own sake.
- Honest waits. If something genuinely takes a minute, say so and offer to notify the user rather than pretending it is nearly done. Long jobs belong on an async path, which is the Batch vs Stream decision.
Bad to Good to Great
Bad - one latency dashboard and a faster model
A single average latency chart. When it rises, swap in a faster model and hope. You cannot tell whether TTFT or completion time moved, you have no per-stage attribution, so the reranker or the cold cache that is actually responsible stays invisible, and the model swap silently costs quality you never measured.
Good - streaming, a max-token cap, and a response cache
Real wins, in the right order, and enough to make most features feel acceptable. What is still missing: no per-stage budget, so the next regression is again a guessing game. Per-call timeouts with no request deadline, so the worst case is the sum of every stage. Nothing specific about the tail, and no defined behaviour when a dependency is slow rather than down.
Great - budgeted stages, deadline-aware execution, and percentile SLOs per route
- TTFT and total completion time budgeted and measured separately, per route, with traces giving per-stage timings.
- Percentile SLOs, not averages, with p95 and p99 alerting and attribution by stage.
- Caching at every layer, including the embedding, reranker, and auth lookups teams forget.
- Output length controlled deliberately by route caps and format instructions, with independent stages parallelised and unnecessary hops removed.
- One deadline propagated through the chain, with a faster model and a degraded answer as defined fallbacks rather than accidents, and fallback and truncation rates tracked as product metrics so silent degradation cannot hide.
- Perceived-latency work treated as first-class β streaming, optimistic rendering, progressive disclosure, real progress for long runs.
When to Use
β Invest in latency engineering when:
- A human is waiting on the output, which is nearly every interactive AI feature
- You have a multi-step chain or agent where per-step tails compound into a slow run
- p95 or p99 is far from your median, meaning a real tail problem rather than a speed problem
- Users complain about slowness but nobody can name which stage is responsible
- A latency SLO is part of the contract for the surface you are shipping
β Do not chase latency when:
- The workload is genuinely asynchronous β batch classification, nightly enrichment, offline evals β where throughput and cost matter and seconds do not
- You have not measured the per-stage decomposition yet, since the first optimisation is almost always in a stage you have not looked at
- Quality is the actual complaint and a faster, weaker model would make the real problem worse
- The win would come from a semantic cache whose correctness risk you have not evaluated
Common Interview Questions
Q1: Your LLM feature feels slow. Where do you start?
By splitting the number, then decomposing the request. First, is TTFT the problem or is total completion time the problem, because they have different causes and different fixes β and if output is streamed, TTFT is what the user perceives. Then I get per-stage timings from traces for a real request: network, auth, query embedding, vector search, reranking, prompt assembly, prefill, decode, validation. The dominant stage is frequently not the model call β a reranker scoring too many candidates, a cold embedding cache, or a cross-region hop to the vector store are common. Only once I know the dominant stage do I pick a technique, because optimising anything else is measurable work with unmeasurable benefit.
Q2: What is the difference between TTFT and total completion time, and which should you optimise?
TTFT is everything up to and including prefill, which processes the whole prompt in one parallel pass. Total completion time is dominated by decode, which emits one token per forward pass sequentially, so it scales with output length. For streamed output to a human, optimise TTFT first β once tokens arrive faster than the user reads, extra completion time is close to invisible. For output a program consumes, or any non-streamed response, total time is the only clock that matters. The biggest lever there is output length, so a max-token cap and an instruction to be concise are genuine latency controls rather than style preferences.
Q3: Explain speculative decoding.
It attacks the fact that decode produces one token per expensive pass over the weights. A small draft model proposes several tokens ahead, cheaply. The large model then verifies that whole proposed span in a single forward pass and accepts the longest prefix matching what it would have generated itself, discarding the rest. Because rejected tokens are thrown away, the output distribution is identical to running the large model alone β it is purely a latency win, not a quality trade. Several tokens per verification pass means lower total completion time. It does nothing for TTFT since prefill is unchanged, and the acceptance rate depends on how predictable the text is, so the gain has to be measured on your own traffic.
Q4: Why do p95 and p99 matter more than the average here, and what does that mean for agents?
Because users experience individual requests, not a distribution, and a session of twenty requests will land in the tail. The tail also has distinct causes worth attributing separately: cold caches, retries, unusually long generations, fleet queueing, cold replicas, and slow external tools. In a multi-step agent it compounds β each stepβs tail is another chance for the run to feel slow, so with an assumed six steps each having a 5 percent chance of exceeding budget, roughly a quarter of runs stall somewhere. The responses are to reduce step count, cap the loop, and show real intermediate progress so a long run stays tolerable.
Q5: How do you stop a slow dependency from blowing your latency SLO?
One deadline for the request, propagated down the chain, instead of independent per-call timeouts that sum to a worst case nobody agreed to. Every stage gets a budget and checks the remaining budget before starting. When the budget is at risk before the model call, I degrade deliberately: route to a smaller faster model, or return a cached or reduced answer, rather than letting the request run long and then fail. Retries only happen when the remaining budget can absorb one, and a circuit breaker stops me queueing behind a provider that is already failing. Critically, I count how often the fallback fires β if it is firing on a large share of traffic it is not a safety valve, it is the production path, and it needs an eval of its own.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts