Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 18 min read

Inference Serving - Complete Deep Dive

Stage 7 - Production Lesson 29 of 36

Prerequisites: How LLMs Actually Work, Tokens and Cost Math, Model Selection and Routing Used in: GPU Cost Engineering, Latency Engineering, ChatGPT


What is Inference Serving?

Inference serving is the layer that turns a set of model weights sitting on a GPU into an endpoint many concurrent users can call. It owns scheduling, memory, batching, and admission β€” everything between β€œthe model can produce tokens” and β€œthe model produces tokens for two thousand people at once without falling over.”

The ChatGPT design covers this at the system level: a fleet, a queue, streaming out to clients. This page is the mechanism underneath one node of that fleet. If you are calling a hosted API you never see any of it, but you pay for its consequences in your bill and your p99, and every interviewer asking β€œwhy does output cost more than input” is asking about the first section below.

Real-world analogy: A restaurant kitchen. Reading the whole ticket and prepping every ingredient happens in one burst, in parallel, with all hands busy β€” that is prefill. Then the dish is plated one component at a time in sequence, and the constraint is not how many cooks you have, it is how fast one person can walk to the pass and back β€” that is decode. A kitchen that seats a new table only when every table currently seated has finished dessert is a kitchen wasting most of its capacity. Continuous batching is seating a new table the moment one gets up.


The Two Phases, and Why They Behave Oppositely

Every generation request runs in two distinct phases with opposite performance characteristics. This is the single most important fact on this page, because almost every cost and latency surprise in LLM serving falls out of it.

flowchart LR
    IN[Prompt tokens arrive] --> PRE[Prefill - all prompt tokens processed in parallel]
    PRE --> KVC[KV Cache - keys and values for every token and layer]
    KVC --> DEC[Decode - one token per forward pass]
    DEC --> KVC
    DEC --> OUT[Token streamed to client]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class IN client
    class PRE,DEC service
    class KVC data
    class OUT edge
Β  Prefill Decode
What it does Processes the entire prompt in one shot Generates one token per forward pass
Parallelism All prompt tokens at once Strictly sequential - token N needs token N-1
Bottleneck Compute - the arithmetic units saturate Memory bandwidth - weights must be read for one token of work
Scales with Prompt length Output length
Owns which metric Time to first token Inter-token latency and total completion time
Batching benefit Modest - already compute-saturated Large - spare bandwidth across the batch

Prefill is compute-bound: there is a lot of arithmetic to do and plenty of work to keep the GPU busy. Decode is memory-bandwidth-bound: to produce one token you stream the model’s weights through the chip, and the arithmetic units are mostly idle waiting on memory. That asymmetry has three consequences worth internalising.

  1. Input and output are not the same product. Output tokens cost more and take longer per token than input tokens, in both hosted pricing and your own hardware, because each one is a separate poorly-utilised pass over the weights. See Tokens and Cost Math for the billing side of this.
  2. Batching fixes decode, not prefill. Because decode wastes bandwidth, running many sequences through the same weight read is nearly free extra throughput. This is why batching is the lever in serving.
  3. A long prompt and a long answer break different things. A long prompt inflates TTFT and prefill compute. A long answer inflates total time and occupies a batch slot for longer, which is a capacity problem, not just a latency one.

The KV Cache, Properly

Self-attention at each step needs a key and a value vector for every token before the current one. Without a cache, generating token 500 would recompute keys and values for all 499 predecessors, and doing that at every step makes total work grow quadratically with sequence length. So they are computed once and kept.

The cache converts quadratic recomputation into linear growth in memory. That trade is overwhelmingly worth it, and it relocates the problem: you no longer have a compute blowup, you have a memory blowup.

So the capacity question on a serving GPU is not β€œhow much compute do I have,” it is β€œhow many KV caches fit in the memory left after the weights.” Concurrency is a memory budget. That reframing is the whole basis of GPU Cost Engineering, and it explains why a model that comfortably serves 64 short chats collapses at 8 long-context ones.


Batching - Bad to Good to Great

Bad - one request at a time

Each request gets the GPU to itself. During decode the arithmetic units sit mostly idle waiting on weight reads, so you are paying for a whole accelerator to do a fraction of its possible work. No production system should serve this way, yet it is exactly what a naive β€œwrap the model in a web handler” deployment does.

Good - static batching

Collect N requests, run them through together, return them together. Throughput jumps immediately, because the weight read that produced one token now produces N.

Then you hit the flaw: the batch is only as fast as its longest sequence. A request that finishes in 20 tokens sits in a completed slot doing nothing until a 900-token sibling is done, and nothing new can be admitted until the whole batch retires. Short requests are held hostage by long ones, so you lose both latency (short requests wait for strangers) and throughput (slots idle at the tail of every batch). The waste grows with the variance of your output lengths β€” and output length variance in a chat product is enormous.

Great - continuous or in-flight batching

The scheduler works at the granularity of a single decode step across a set of sequences, not a fixed batch. After each step, finished sequences are evicted and waiting requests are admitted into the freed slots immediately.

flowchart TD
    subgraph Static["Static batching"]
        S1[Batch admitted together] --> S2[Short sequence finishes early]
        S2 --> S3[Slot sits idle until the longest sequence finishes]
        S3 --> S4[Entire batch returns and only then is a new batch admitted]
    end

    subgraph Continuous["Continuous or in-flight batching"]
        C1[Running set of sequences] --> C2[One decode step for every active sequence]
        C2 --> C3[Finished sequences evicted and their memory freed]
        C3 --> C4[Scheduler admits waiting requests into the freed slots]
        C4 --> C1
    end

    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class S1,S2,S3,S4 async
    class C1,C2 service
    class C3,C4 data

A sequence is now independent: it arrives, joins the running set, and leaves when it emits its stop token. No request waits on an unrelated request’s length. This is the single largest throughput win available in self-hosted serving, and it is the default in every modern stack. Long prompts can still stall decode for everyone while a big prefill runs, which is why stacks add chunked prefill β€” slicing a long prefill across several scheduler steps so in-flight decodes keep progressing.


PagedAttention - Virtual Memory for the KV Cache

Continuous batching exposes a memory problem. If each sequence’s KV cache must be one contiguous allocation, you have to reserve space for its maximum possible length up front, because you cannot know how long the answer will be. Two kinds of waste follow: over-reservation (a request that reserved room for 4,000 tokens and emitted 40 wasted the rest) and fragmentation (freed blocks of the wrong size and shape cannot be reused).

PagedAttention borrows the operating system’s answer. Instead of one contiguous region per sequence, the KV cache is carved into fixed-size blocks, and each sequence holds a block table mapping its logical token positions to physical blocks anywhere in memory. Attention kernels are written to follow the block table.

The payoff is not elegance, it is batch size. Reclaimed memory becomes more concurrent sequences, and more concurrent sequences is more throughput per GPU. This is why PagedAttention and continuous batching are always discussed together: the scheduler creates the opportunity and the memory manager is what lets you take it. The same β€œindirection makes a fixed resource go further” idea appears throughout Caching.


Prefix Sharing

If a thousand requests begin with the same 800-token system prompt, prefill is recomputing identical keys and values a thousand times and storing a thousand identical copies. With block-based KV storage you can instead compute the prefix once and let every request point at the same blocks, copying only when a sequence diverges.

Two wins, and it is worth keeping them separate because they are often conflated:

It only works on an exact prefix match from position zero, which is a concrete design constraint on how you assemble prompts: put the stable material β€” system prompt, tool definitions, few-shot examples β€” first, and the volatile material β€” retrieved chunks, the user’s question, a timestamp β€” last. A per-request session ID injected at the top of the prompt silently destroys sharing for everyone. Hosted providers expose the same mechanism as prompt caching; see Prompt and Semantic Caching.


Throughput Versus Latency Is a Choice You Make

Raising the batch size raises tokens per second per GPU and raises per-request latency. Both, always, together. More sequences share each weight read, so aggregate output climbs; but each individual sequence now waits behind more work per step, so its inter-token latency stretches.

There is no setting that wins both. What you have is a dial between cost per token and responsiveness, and the only principled way to set it is from an SLO:

  1. Write down the latency target that matters β€” usually a p95 on time to first token and on inter-token latency.
  2. Load-test upward through batch sizes, measuring throughput and the latency percentiles together.
  3. Take the largest batch size that still meets the SLO. That is your setting; anything larger is cheaper tokens your users will feel.

Two fleets beat one compromise. An interactive fleet runs a modest batch size and tight queue limits; a batch fleet runs the largest batch that fits memory and does not care about p99. Mixing a nightly bulk classification job into your chat fleet degrades chat latency to buy throughput nobody asked for. Percentile discipline generally is in Performance Metrics.


Quantization and Serving Capacity

Quantization stores weights β€” and optionally the KV cache β€” at lower numeric precision. Its effect on serving is mostly a memory effect, and memory is the binding constraint, so the consequences cascade:

The cost is quality, and it is task-dependent and not always visible in casual testing. Treat any quantized deployment as a new model: run it through your eval suite and compare against the full-precision baseline on your task before believing the throughput win is free. Never quote a quality-neutral claim you have not measured.


Speculative Decoding, in One Paragraph

Decode is sequential and bandwidth-bound, so the obvious attack is to get more than one token per expensive pass. A small draft model proposes several tokens ahead; the large model verifies them in a single forward pass and accepts the longest correct run. Output is identical to what the large model would have produced alone, because rejected tokens are discarded. It reduces decode latency, does nothing for TTFT, and its benefit depends heavily on how predictable your text is. Full treatment on Latency Engineering.


Technique Scorecard

Technique What it improves What it costs
Continuous batching Throughput and queue wait, by a large margin Scheduler complexity; per-token latency varies with fleet load
Chunked prefill Stops long prompts stalling in-flight decodes Slight prefill overhead and more scheduling work
PagedAttention Memory efficiency, which converts into larger batches Custom attention kernels; couples you to the serving stack
Prefix sharing TTFT and prefill compute on shared system prompts Needs exact prefix match; cache memory and an eviction policy
Weight quantization Memory headroom, decode bandwidth, sometimes GPU count Quality risk that must be eval-gated per task
KV cache quantization Concurrency at a fixed memory budget Quality risk on long contexts; stack support varies
Speculative decoding Decode latency and total completion time Draft model memory and compute; gains are workload-dependent
Tensor parallelism Fits larger models; can cut decode latency Inter-GPU communication; needs a fast interconnect
Larger max batch size Tokens per second per GPU Per-request latency, directly
Queueing and admission control Predictable tails under overload Rejected or delayed work; requires honest capacity planning

The Operational Surface

The serving algorithm is the interesting part. The operational reality is what pages you.

Model loading is slow. Weights have to be fetched and moved onto the device before a replica serves anything, and for large models that is minutes, not seconds. Every autoscaling decision inherits that number.

Warmup is mandatory. A freshly loaded replica has cold kernels, unbuilt graphs, and possibly a just-in-time compilation step. Send it synthetic requests until latency settles, and do not let the load balancer route real traffic to a replica that has not passed a warm readiness check.

Autoscaling a GPU fleet is not autoscaling a stateless web tier. Instances are scarce, expensive, and slow to come up, so reactive scaling on a traffic spike arrives after the spike. What works: scale on a leading indicator such as queue depth or tokens in flight rather than on CPU; keep warm headroom sized to your load time; scale in slowly and with generous cooldowns so you are not paying the load cost twice; and pre-scale on known schedules where traffic is predictable.

Queueing plus admission control beats unbounded concurrency. If you accept every request, an overloaded replica admits sequences until KV memory is exhausted, then starts preempting or swapping in-flight sequences, and every latency percentile degrades at once β€” including for the requests that were already nearly done. The alternative is boring and correct: an explicit bounded queue, a cap on concurrent sequences derived from measured KV capacity, a cap on max tokens per request, and a fast rejection or a clear retry signal when the queue is full. Shedding 2% of load protects the other 98%; admitting everything degrades all of it. Pair it with Rate Limiting at the edge and Circuit Breaker on the caller side.

Know what your stack is for. Vendor-neutrally: vLLM, TensorRT-LLM, SGLang, and TGI are the mainstream GPU serving stacks, and llama.cpp targets small-scale and on-device deployment with aggressive quantization. They differ in scheduler behaviour, quantization support, structured-output support, and how much hardware-specific tuning they expect. Benchmark two on your own model and your own traffic shape rather than trusting anyone’s published numbers, including theirs.


When to Use

βœ… Self-host and tune a serving stack when:

❌ Stay on a hosted API when:


Common Interview Questions

Q1: Why do output tokens cost more than input tokens?

Because they are produced by a different phase with a different bottleneck. Prefill processes the whole prompt in one parallel pass and saturates the arithmetic units, so input tokens are cheap per token. Decode produces exactly one token per forward pass, and each pass streams the model weights through the chip, so it is memory-bandwidth-bound and badly utilised. You are paying for a full pass over the weights per output token, and that pass cannot be parallelised within a request because token N depends on token N-1. Hosted pricing reflects the same physics that shows up in your own throughput numbers when you self-host.

Q2: What actually limits how many concurrent requests one GPU can serve?

KV cache memory, not compute. The weights take a fixed slice of device memory; what is left has to hold one KV cache per in-flight sequence, and that cache grows with every token generated, for every layer and head. So concurrency is a memory budget, and it is a function of context length β€” the same GPU might comfortably run a large batch of short chats and fall over on a handful of long-context requests. That is why PagedAttention matters: eliminating over-reservation and fragmentation in KV storage buys you batch size, and batch size is throughput.

Q3: Explain continuous batching and what problem it solves.

Static batching admits a fixed group of requests, runs them together, and returns them together, so the batch runs at the pace of its longest sequence. A 20-token answer occupies a slot until a 900-token sibling finishes, and no new work is admitted until the whole batch retires β€” you lose latency on short requests and throughput on idle tail slots, and the loss scales with output-length variance, which in a chat product is huge. Continuous batching schedules one decode step across the currently active sequences, evicts each sequence the moment it finishes, and admits waiting requests into the freed slots immediately. Requests become independent of each other’s length.

Q4: How do you choose a batch size?

From the SLO, by measurement. Larger batches always raise tokens per second per GPU and always raise per-request latency, so there is no optimum β€” there is a dial between cost per token and responsiveness. I write down the p95 targets for time to first token and inter-token latency, load-test upward through batch sizes while recording throughput and percentiles together, and take the largest batch that still meets the targets. Then I separate fleets: interactive traffic gets a modest batch and a bounded queue, bulk work gets the largest batch memory allows and no latency promise.

Q5: Your GPU fleet is saturated and p99 has collapsed. What do you do?

First stop the collapse, then fix the cause. The immediate move is admission control β€” a bounded queue, a cap on concurrent sequences derived from measured KV capacity, and a max-token cap per request β€” because an unbounded fleet under overload thrashes KV memory and degrades every request including the nearly-finished ones. Shedding a small fraction of load protects the rest. Then diagnose with the two phases in mind: if TTFT is what degraded, prompts or prefill queueing are the problem, so look at context length and prefix sharing; if inter-token latency degraded, the batch is too large or too many long generations are in flight, so look at max-token caps and fleet split. Only after that do I add GPUs, since capacity bought to cover an unbounded concurrency bug just gets consumed the same way.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access