Inference Serving - Complete Deep Dive
Prerequisites: How LLMs Actually Work, Tokens and Cost Math, Model Selection and Routing Used in: GPU Cost Engineering, Latency Engineering, ChatGPT
What is Inference Serving?
Inference serving is the layer that turns a set of model weights sitting on a GPU into an endpoint many concurrent users can call. It owns scheduling, memory, batching, and admission β everything between βthe model can produce tokensβ and βthe model produces tokens for two thousand people at once without falling over.β
The ChatGPT design covers this at the system level: a fleet, a queue, streaming out to clients. This page is the mechanism underneath one node of that fleet. If you are calling a hosted API you never see any of it, but you pay for its consequences in your bill and your p99, and every interviewer asking βwhy does output cost more than inputβ is asking about the first section below.
Real-world analogy: A restaurant kitchen. Reading the whole ticket and prepping every ingredient happens in one burst, in parallel, with all hands busy β that is prefill. Then the dish is plated one component at a time in sequence, and the constraint is not how many cooks you have, it is how fast one person can walk to the pass and back β that is decode. A kitchen that seats a new table only when every table currently seated has finished dessert is a kitchen wasting most of its capacity. Continuous batching is seating a new table the moment one gets up.
The Two Phases, and Why They Behave Oppositely
Every generation request runs in two distinct phases with opposite performance characteristics. This is the single most important fact on this page, because almost every cost and latency surprise in LLM serving falls out of it.
flowchart LR
IN[Prompt tokens arrive] --> PRE[Prefill - all prompt tokens processed in parallel]
PRE --> KVC[KV Cache - keys and values for every token and layer]
KVC --> DEC[Decode - one token per forward pass]
DEC --> KVC
DEC --> OUT[Token streamed to client]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class IN client
class PRE,DEC service
class KVC data
class OUT edge
| Β | Prefill | Decode |
|---|---|---|
| What it does | Processes the entire prompt in one shot | Generates one token per forward pass |
| Parallelism | All prompt tokens at once | Strictly sequential - token N needs token N-1 |
| Bottleneck | Compute - the arithmetic units saturate | Memory bandwidth - weights must be read for one token of work |
| Scales with | Prompt length | Output length |
| Owns which metric | Time to first token | Inter-token latency and total completion time |
| Batching benefit | Modest - already compute-saturated | Large - spare bandwidth across the batch |
Prefill is compute-bound: there is a lot of arithmetic to do and plenty of work to keep the GPU busy. Decode is memory-bandwidth-bound: to produce one token you stream the modelβs weights through the chip, and the arithmetic units are mostly idle waiting on memory. That asymmetry has three consequences worth internalising.
- Input and output are not the same product. Output tokens cost more and take longer per token than input tokens, in both hosted pricing and your own hardware, because each one is a separate poorly-utilised pass over the weights. See Tokens and Cost Math for the billing side of this.
- Batching fixes decode, not prefill. Because decode wastes bandwidth, running many sequences through the same weight read is nearly free extra throughput. This is why batching is the lever in serving.
- A long prompt and a long answer break different things. A long prompt inflates TTFT and prefill compute. A long answer inflates total time and occupies a batch slot for longer, which is a capacity problem, not just a latency one.
The KV Cache, Properly
Self-attention at each step needs a key and a value vector for every token before the current one. Without a cache, generating token 500 would recompute keys and values for all 499 predecessors, and doing that at every step makes total work grow quadratically with sequence length. So they are computed once and kept.
The cache converts quadratic recomputation into linear growth in memory. That trade is overwhelmingly worth it, and it relocates the problem: you no longer have a compute blowup, you have a memory blowup.
- The cache holds keys and values for every token, every layer, every attention head, per sequence.
- It grows by one tokenβs worth on every decode step, for the whole life of the request.
- It is per sequence β nothing about it is shareable between two different conversations, except a shared prefix (below).
- A GPU serving a model holds weights plus one KV cache per concurrent sequence. Weights are fixed; KV cache is not.
So the capacity question on a serving GPU is not βhow much compute do I have,β it is βhow many KV caches fit in the memory left after the weights.β Concurrency is a memory budget. That reframing is the whole basis of GPU Cost Engineering, and it explains why a model that comfortably serves 64 short chats collapses at 8 long-context ones.
Batching - Bad to Good to Great
Bad - one request at a time
Each request gets the GPU to itself. During decode the arithmetic units sit mostly idle waiting on weight reads, so you are paying for a whole accelerator to do a fraction of its possible work. No production system should serve this way, yet it is exactly what a naive βwrap the model in a web handlerβ deployment does.
Good - static batching
Collect N requests, run them through together, return them together. Throughput jumps immediately, because the weight read that produced one token now produces N.
Then you hit the flaw: the batch is only as fast as its longest sequence. A request that finishes in 20 tokens sits in a completed slot doing nothing until a 900-token sibling is done, and nothing new can be admitted until the whole batch retires. Short requests are held hostage by long ones, so you lose both latency (short requests wait for strangers) and throughput (slots idle at the tail of every batch). The waste grows with the variance of your output lengths β and output length variance in a chat product is enormous.
Great - continuous or in-flight batching
The scheduler works at the granularity of a single decode step across a set of sequences, not a fixed batch. After each step, finished sequences are evicted and waiting requests are admitted into the freed slots immediately.
flowchart TD
subgraph Static["Static batching"]
S1[Batch admitted together] --> S2[Short sequence finishes early]
S2 --> S3[Slot sits idle until the longest sequence finishes]
S3 --> S4[Entire batch returns and only then is a new batch admitted]
end
subgraph Continuous["Continuous or in-flight batching"]
C1[Running set of sequences] --> C2[One decode step for every active sequence]
C2 --> C3[Finished sequences evicted and their memory freed]
C3 --> C4[Scheduler admits waiting requests into the freed slots]
C4 --> C1
end
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class S1,S2,S3,S4 async
class C1,C2 service
class C3,C4 data
A sequence is now independent: it arrives, joins the running set, and leaves when it emits its stop token. No request waits on an unrelated requestβs length. This is the single largest throughput win available in self-hosted serving, and it is the default in every modern stack. Long prompts can still stall decode for everyone while a big prefill runs, which is why stacks add chunked prefill β slicing a long prefill across several scheduler steps so in-flight decodes keep progressing.
PagedAttention - Virtual Memory for the KV Cache
Continuous batching exposes a memory problem. If each sequenceβs KV cache must be one contiguous allocation, you have to reserve space for its maximum possible length up front, because you cannot know how long the answer will be. Two kinds of waste follow: over-reservation (a request that reserved room for 4,000 tokens and emitted 40 wasted the rest) and fragmentation (freed blocks of the wrong size and shape cannot be reused).
PagedAttention borrows the operating systemβs answer. Instead of one contiguous region per sequence, the KV cache is carved into fixed-size blocks, and each sequence holds a block table mapping its logical token positions to physical blocks anywhere in memory. Attention kernels are written to follow the block table.
- Blocks are allocated as the sequence grows, so nothing is reserved for tokens that never get generated.
- Any freed block fits any sequence, so external fragmentation largely disappears.
- Waste is bounded by the partially-filled last block per sequence rather than by the gap between reserved and actual length.
The payoff is not elegance, it is batch size. Reclaimed memory becomes more concurrent sequences, and more concurrent sequences is more throughput per GPU. This is why PagedAttention and continuous batching are always discussed together: the scheduler creates the opportunity and the memory manager is what lets you take it. The same βindirection makes a fixed resource go furtherβ idea appears throughout Caching.
Prefix Sharing
If a thousand requests begin with the same 800-token system prompt, prefill is recomputing identical keys and values a thousand times and storing a thousand identical copies. With block-based KV storage you can instead compute the prefix once and let every request point at the same blocks, copying only when a sequence diverges.
Two wins, and it is worth keeping them separate because they are often conflated:
- Compute: the shared prefix is prefilled once, not per request, which cuts TTFT directly.
- Memory: one copy of the prefix cache instead of N, which returns capacity to the batch.
It only works on an exact prefix match from position zero, which is a concrete design constraint on how you assemble prompts: put the stable material β system prompt, tool definitions, few-shot examples β first, and the volatile material β retrieved chunks, the userβs question, a timestamp β last. A per-request session ID injected at the top of the prompt silently destroys sharing for everyone. Hosted providers expose the same mechanism as prompt caching; see Prompt and Semantic Caching.
Throughput Versus Latency Is a Choice You Make
Raising the batch size raises tokens per second per GPU and raises per-request latency. Both, always, together. More sequences share each weight read, so aggregate output climbs; but each individual sequence now waits behind more work per step, so its inter-token latency stretches.
There is no setting that wins both. What you have is a dial between cost per token and responsiveness, and the only principled way to set it is from an SLO:
- Write down the latency target that matters β usually a p95 on time to first token and on inter-token latency.
- Load-test upward through batch sizes, measuring throughput and the latency percentiles together.
- Take the largest batch size that still meets the SLO. That is your setting; anything larger is cheaper tokens your users will feel.
Two fleets beat one compromise. An interactive fleet runs a modest batch size and tight queue limits; a batch fleet runs the largest batch that fits memory and does not care about p99. Mixing a nightly bulk classification job into your chat fleet degrades chat latency to buy throughput nobody asked for. Percentile discipline generally is in Performance Metrics.
Quantization and Serving Capacity
Quantization stores weights β and optionally the KV cache β at lower numeric precision. Its effect on serving is mostly a memory effect, and memory is the binding constraint, so the consequences cascade:
- Smaller weights leave more room for KV cache, which means a larger batch, which means more throughput per GPU.
- Less data to stream per forward pass helps the bandwidth-bound decode phase directly.
- A model that previously needed multiple GPUs may fit on one, removing inter-GPU communication from the decode path.
- Quantizing the KV cache itself is a second, independent lever, and it buys concurrency rather than weight space.
The cost is quality, and it is task-dependent and not always visible in casual testing. Treat any quantized deployment as a new model: run it through your eval suite and compare against the full-precision baseline on your task before believing the throughput win is free. Never quote a quality-neutral claim you have not measured.
Speculative Decoding, in One Paragraph
Decode is sequential and bandwidth-bound, so the obvious attack is to get more than one token per expensive pass. A small draft model proposes several tokens ahead; the large model verifies them in a single forward pass and accepts the longest correct run. Output is identical to what the large model would have produced alone, because rejected tokens are discarded. It reduces decode latency, does nothing for TTFT, and its benefit depends heavily on how predictable your text is. Full treatment on Latency Engineering.
Technique Scorecard
| Technique | What it improves | What it costs |
|---|---|---|
| Continuous batching | Throughput and queue wait, by a large margin | Scheduler complexity; per-token latency varies with fleet load |
| Chunked prefill | Stops long prompts stalling in-flight decodes | Slight prefill overhead and more scheduling work |
| PagedAttention | Memory efficiency, which converts into larger batches | Custom attention kernels; couples you to the serving stack |
| Prefix sharing | TTFT and prefill compute on shared system prompts | Needs exact prefix match; cache memory and an eviction policy |
| Weight quantization | Memory headroom, decode bandwidth, sometimes GPU count | Quality risk that must be eval-gated per task |
| KV cache quantization | Concurrency at a fixed memory budget | Quality risk on long contexts; stack support varies |
| Speculative decoding | Decode latency and total completion time | Draft model memory and compute; gains are workload-dependent |
| Tensor parallelism | Fits larger models; can cut decode latency | Inter-GPU communication; needs a fast interconnect |
| Larger max batch size | Tokens per second per GPU | Per-request latency, directly |
| Queueing and admission control | Predictable tails under overload | Rejected or delayed work; requires honest capacity planning |
The Operational Surface
The serving algorithm is the interesting part. The operational reality is what pages you.
Model loading is slow. Weights have to be fetched and moved onto the device before a replica serves anything, and for large models that is minutes, not seconds. Every autoscaling decision inherits that number.
Warmup is mandatory. A freshly loaded replica has cold kernels, unbuilt graphs, and possibly a just-in-time compilation step. Send it synthetic requests until latency settles, and do not let the load balancer route real traffic to a replica that has not passed a warm readiness check.
Autoscaling a GPU fleet is not autoscaling a stateless web tier. Instances are scarce, expensive, and slow to come up, so reactive scaling on a traffic spike arrives after the spike. What works: scale on a leading indicator such as queue depth or tokens in flight rather than on CPU; keep warm headroom sized to your load time; scale in slowly and with generous cooldowns so you are not paying the load cost twice; and pre-scale on known schedules where traffic is predictable.
Queueing plus admission control beats unbounded concurrency. If you accept every request, an overloaded replica admits sequences until KV memory is exhausted, then starts preempting or swapping in-flight sequences, and every latency percentile degrades at once β including for the requests that were already nearly done. The alternative is boring and correct: an explicit bounded queue, a cap on concurrent sequences derived from measured KV capacity, a cap on max tokens per request, and a fast rejection or a clear retry signal when the queue is full. Shedding 2% of load protects the other 98%; admitting everything degrades all of it. Pair it with Rate Limiting at the edge and Circuit Breaker on the caller side.
Know what your stack is for. Vendor-neutrally: vLLM, TensorRT-LLM, SGLang, and TGI are the mainstream GPU serving stacks, and llama.cpp targets small-scale and on-device deployment with aggressive quantization. They differ in scheduler behaviour, quantization support, structured-output support, and how much hardware-specific tuning they expect. Benchmark two on your own model and your own traffic shape rather than trusting anyoneβs published numbers, including theirs.
When to Use
β Self-host and tune a serving stack when:
- Volume is high and steady enough that fixed GPU cost beats per-token API pricing β check the method in GPU Cost Engineering
- You need open-weights models, a fine-tuned checkpoint, or full control over the model version
- Data residency or privacy rules prevent sending traffic to a hosted provider
- You need serving-level control such as prefix sharing, custom sampling, or guaranteed isolation
β Stay on a hosted API when:
- Traffic is spiky or low, so owned GPUs would sit idle most of the day
- Nobody on the team wants to own GPU capacity, warmup, and autoscaling
- You are still finding product-market fit and the model choice may change next month
- Your quality bar needs a frontier model you cannot self-host anyway
Common Interview Questions
Q1: Why do output tokens cost more than input tokens?
Because they are produced by a different phase with a different bottleneck. Prefill processes the whole prompt in one parallel pass and saturates the arithmetic units, so input tokens are cheap per token. Decode produces exactly one token per forward pass, and each pass streams the model weights through the chip, so it is memory-bandwidth-bound and badly utilised. You are paying for a full pass over the weights per output token, and that pass cannot be parallelised within a request because token N depends on token N-1. Hosted pricing reflects the same physics that shows up in your own throughput numbers when you self-host.
Q2: What actually limits how many concurrent requests one GPU can serve?
KV cache memory, not compute. The weights take a fixed slice of device memory; what is left has to hold one KV cache per in-flight sequence, and that cache grows with every token generated, for every layer and head. So concurrency is a memory budget, and it is a function of context length β the same GPU might comfortably run a large batch of short chats and fall over on a handful of long-context requests. That is why PagedAttention matters: eliminating over-reservation and fragmentation in KV storage buys you batch size, and batch size is throughput.
Q3: Explain continuous batching and what problem it solves.
Static batching admits a fixed group of requests, runs them together, and returns them together, so the batch runs at the pace of its longest sequence. A 20-token answer occupies a slot until a 900-token sibling finishes, and no new work is admitted until the whole batch retires β you lose latency on short requests and throughput on idle tail slots, and the loss scales with output-length variance, which in a chat product is huge. Continuous batching schedules one decode step across the currently active sequences, evicts each sequence the moment it finishes, and admits waiting requests into the freed slots immediately. Requests become independent of each otherβs length.
Q4: How do you choose a batch size?
From the SLO, by measurement. Larger batches always raise tokens per second per GPU and always raise per-request latency, so there is no optimum β there is a dial between cost per token and responsiveness. I write down the p95 targets for time to first token and inter-token latency, load-test upward through batch sizes while recording throughput and percentiles together, and take the largest batch that still meets the targets. Then I separate fleets: interactive traffic gets a modest batch and a bounded queue, bulk work gets the largest batch memory allows and no latency promise.
Q5: Your GPU fleet is saturated and p99 has collapsed. What do you do?
First stop the collapse, then fix the cause. The immediate move is admission control β a bounded queue, a cap on concurrent sequences derived from measured KV capacity, and a max-token cap per request β because an unbounded fleet under overload thrashes KV memory and degrades every request including the nearly-finished ones. Shedding a small fraction of load protects the rest. Then diagnose with the two phases in mind: if TTFT is what degraded, prompts or prefill queueing are the problem, so look at context length and prefix sharing; if inter-token latency degraded, the batch is too large or too many long generations are in flight, so look at max-token caps and fleet split. Only after that do I add GPUs, since capacity bought to cover an unbounded concurrency bug just gets consumed the same way.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts