GPU Capacity and Cost Engineering - Complete Deep Dive
Prerequisites: Inference Serving, Tokens and Cost Math, Back-of-Envelope Estimation Used in: Latency Engineering, Model Selection and Routing, ChatGPT
What is GPU Cost Engineering?
GPU cost engineering is the practice of deciding how much accelerator capacity you need, whether you should own it at all, and which levers actually reduce the bill โ with measurements rather than intuition. It sits directly on top of Inference Serving: the serving stack determines how much work one GPU can do, and this page turns that number into a fleet size and a cost per million tokens.
Real-world analogy: Taxis versus a company car. A taxi charges per trip โ no trips, no bill, and the per-trip price never improves no matter how much you ride. A company car costs the same every month whether it is on the motorway or in the car park, so it only wins if you drive it constantly. Nobody buys a car for four trips a month, and nobody takes taxis for a 300-kilometre daily commute. Hosted API versus owned GPUs is exactly this decision, and utilisation is the only thing that decides it.
The Economics Are Structurally Different
flowchart TD
subgraph Hosted["Hosted API - variable cost"]
H1[Request arrives] --> H2[Pay per token consumed]
H2 --> H3[No traffic means no spend]
H3 --> H4[Cost tracks demand exactly and never amortises]
end
subgraph Owned["Self hosted GPUs - fixed cost"]
O1[Provision capacity for peak] --> O2[Pay per GPU hour idle or saturated]
O2 --> O3[No traffic still costs full price]
O3 --> O4[Cost per token falls only as utilisation rises]
end
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
classDef async fill:#b4f,stroke:#333,color:#000
classDef client fill:#f97316,stroke:#c2410c,color:#fff
class H1 client
class H2,H3,H4 service
class O1 client
class O2,O3,O4 data
A hosted API is variable cost per token. It is linear, it has no floor, and it has no volume-driven improvement unless you negotiate one. A self-hosted GPU is fixed cost per hour โ the meter runs identically at 4% utilisation and at 94%.
Three consequences that decide most architecture arguments:
- Utilisation is the entire game. Cost per token on owned capacity is fleet cost divided by delivered tokens. The numerator is fixed by your provisioning decision; the denominator is whatever traffic shows up. Double utilisation and you have halved cost per token without touching the model.
- Spiky traffic is the worst possible fit for owned capacity. You must provision for peak but you pay for the troughs, so a workload with a 10x peak-to-trough ratio is paying roughly 10x its marginal cost during the quiet hours. Business-hours-only internal tools are the classic mismatch: sized for 9am, idle for sixteen hours a day.
- Fixed cost punishes uncertainty. Hosted spend scales down with a failed launch. A reserved GPU fleet does not, and a one-year commitment on a product that might get cancelled in three months is a bet on the roadmap, not on the infrastructure.
This is also why the hybrid answer is so common and so sensible: own capacity for the steady base load, burst to a hosted API above it. The economics of each shape are matched to the traffic that suits them.
Memory Is the Binding Constraint
Before sizing anything, internalise what fills up on a serving GPU. Device memory holds three things:
- Model weights โ fixed, paid once per replica, independent of traffic.
- KV cache, per concurrent sequence โ grows with every token generated, for every layer and head, and is released only when the sequence finishes.
- Activation and framework overhead โ working memory for the forward pass, plus whatever the runtime reserves.
Weights and overhead are constants. The KV cache is the variable, and it is the one that runs out. The consequence is the sentence to remember: concurrency is limited by KV cache memory, not by compute. You do not run out of arithmetic; you run out of room to remember what has already been said.
That is why context length is a capacity decision and not just a cost decision. Doubling average context roughly doubles KV cache per sequence, which roughly halves how many sequences fit, which halves your batch size and therefore your throughput per GPU. A feature that quietly grew its retrieved context from four chunks to twelve did not just raise the token bill โ it shrank the fleetโs capacity, and that shows up as queueing and p99 rather than as a line item.
Worked Sizing Example
Every number in this section is an ASSUMPTION invented to make the arithmetic readable. None of it is a quote, a benchmark, or a spec for any real accelerator. Real serving throughput and real KV capacity must be measured on your own model, your own context lengths, and your own hardware, and hourly rates must be looked up on your providerโs current pricing page. What transfers is the method and the shape of the result, never the figures.
The hourly rate appears below as R โ one GPU-hour, as a unit. Substitute your real rate at the end and the whole calculation scales linearly.
Step 1 - demand, from the SLO
- ASSUMED peak concurrent active generations: 400
- ASSUMED per-sequence output rate needed for a comfortable stream: 25 output tokens per second
- Aggregate peak demand: 400 ร 25 = 10,000 output tokens per second
Note that the SLO is what converts โ400 usersโ into a token rate. Accept a slower stream and the same users need less capacity โ the dial from Inference Serving, applied to a budget.
Step 2 - supply, from a measurement
- ASSUMED measured sustained aggregate throughput of one GPU, at the largest batch size that still met the latency SLO: 2,000 output tokens per second
- Throughput-derived GPU count: 10,000 รท 2,000 = 5
- ASSUMED share of capacity consumed by prefill rather than decode: 20%, so 5 รท 0.8 = 6.25 โ 7
Step 3 - the memory check, which usually wins
- ASSUMED measured concurrent sequences one GPU holds at your average context length before KV memory is exhausted: 48
- Memory-derived GPU count: 400 รท 48 = 8.3 โ 9
Size on the larger of the two. Nine GPUs, because memory binds before compute does โ which is the normal outcome, and the reason a purely throughput-based estimate under-provisions. Add one replica for failure headroom and enough warm headroom to cover model load time, and you are provisioning 10 for a peak of 400.
Step 4 - cost per million tokens, where utilisation appears
- Fleet cost: 9 ร R per hour โ 9 ร 730 = 6,570 R per month
- Capacity at full tilt: 9 ร 2,000 = 18,000 tokens per second โ about 47,300 million output tokens per month
| ASSUMED average utilisation | Delivered output tokens per month | Cost per million output tokens |
|---|---|---|
| 100% - unattainable | 47,300 M | 0.14 R |
| 40% | 18,900 M | 0.35 R |
| 10% | 4,730 M | 1.39 R |
The ratio is the lesson. Cost per token is linear in utilisation, so a 10x utilisation gap is a 10x cost gap on identical hardware serving an identical model. No quantization scheme, batch size, or kernel optimisation will recover a 10x. Filling the fleet will.
The general discipline โ assume loudly, compute roughly, check the order of magnitude โ is in Back-of-Envelope Estimation.
Provisioning - Bad to Good to Great
Bad - provision for peak and leave it
Size the fleet for the worst hour of the week, run it flat out at that size forever, and treat the bill as the cost of doing business. This is the most common owned-capacity deployment and it is structurally wasteful: average utilisation ends up a fraction of peak, so most of what you pay for is idle silicon. It also fails in the other direction, because โpeakโ was measured before the feature got popular.
Good - autoscale the fleet on demand
Scale replicas up and down with traffic. A real improvement, and it removes the overnight waste. Its limits are specific to accelerators: instances are scarce and slow to start, so reactive scaling arrives after the spike it was meant to absorb; warm headroom has to be paid for anyway to cover model load time; and scaling in aggressively means paying the load cost repeatedly when traffic oscillates. You also still have troughs โ smaller ones, but empty.
Great - a two-class fleet with queued backfill work
Provision the steady base load as owned capacity you keep busy, burst above it to a hosted API rather than to cold GPUs, and run a low-priority batch class that consumes whatever the interactive class is not using. Interactive traffic gets a modest batch size, a bounded queue, and preemption rights. Batch work โ evals, embedding backfills, bulk classification โ gets the largest batch memory allows, spot or preemptible capacity where available, and no latency promise. Utilisation rises toward what the hardware can actually sustain, and the peak is absorbed by the cost structure that handles peaks well.
The Crossover Analysis - A Method, Not a Verdict
The honest answer to โis self-hosting cheaperโ is โcompute it for your volume and your utilisation, and recompute it when either changes.โ What follows is the procedure, not a conclusion.
- Measure your actual token volume, split into input and output, per feature. If you cannot produce this number you are not ready to have the conversation. Instrument it as described in Tokens and Cost Math.
- Price the hosted path at current published rates for the model you would actually ship, using your real input-to-output split. Include retries, eval traffic, and internal testing โ teams routinely forget all three.
- Price the self-hosted path using the sizing above: fleet size from the memory-bound constraint, hourly rate from your provider, then divide by delivered tokens at your realistic utilisation, not at capacity.
- Add the costs that are not on the GPU invoice. Engineer time to operate the fleet, on-call burden, model loading and warmup waste, over-provisioned headroom, load testing, and the opportunity cost of the work those engineers are not doing. This is frequently larger than the hardware delta and is almost always left out of the spreadsheet that โprovesโ self-hosting wins.
- Plot both against volume and find where the lines cross. Then ask the real question: how confident are you that you will sustain the volume and utilisation on the right-hand side of that crossing?
Two things are usually true. Self-hosting can be dramatically cheaper per token at high, steady, well-utilised volume with an open-weights model that is good enough for the task. And most teams do not have high, steady, well-utilised volume โ they have a launch curve, a weekday shape, and a model choice that is still moving. Break-even assumes the utilisation you will actually achieve, not the utilisation your capacity permits.
A useful middle path: prove the workload on a hosted API, measure the real traffic shape for a quarter, and only then move the steady base load onto owned capacity while bursting the peaks. You get the cost structure that fits each part of the curve. The model-side version of this decision is in Model Selection and Routing.
The Levers, Ranked
- Route easy traffic to a smaller model. The spread between model tiers is the largest single factor in most bills, and most production traffic is not hard. Classification, routing, extraction, and short factual answers rarely need your best model. Gate it with evals so you find out when the router is wrong. See Model Selection and Routing and Distillation and Small Models.
- Cut context length. Applies to every single request, needs no new infrastructure, and on owned capacity it pays twice โ less prefill compute and a smaller KV cache per sequence, which means a larger batch and more throughput per GPU. Retrieve three excellent chunks instead of twelve mediocre ones.
- Cache. Repeated traffic is free traffic. Exact-match caching is trivial, prefix and prompt caching discount a stable system prompt, and semantic caching catches near-duplicates at the price of a threshold you must tune. See Prompt and Semantic Caching and the mechanics in Caching.
- Quantize. On owned capacity this is a memory lever that converts into a throughput lever. Eval-gate it; a quality regression you did not measure is not a saving.
- Raise batch size to the edge of the SLO. Free throughput up to the point where p95 latency breaches the target, and not one step further.
- Use spot or preemptible capacity for batch work. Meaningful discounts for work that can be interrupted and retried. Never for the interactive path.
- Separate interactive from batch fleets. Mostly an enabler for the two levers above, and it stops a nightly bulk job from degrading chat latency.
- Cap maximum output tokens. A ceiling on the most expensive phase, and simultaneously a latency control.
| Lever | Typical magnitude | Tradeoff |
|---|---|---|
| Route to a smaller model | Largest available - often a multiple, not a percentage | Router complexity; misroutes are quality regressions unless eval-gated |
| Cut context length | Large, and it applies to every request | Retrieval tuning effort; too aggressive and answers lose grounding |
| Caching | Large where traffic repeats, negligible where it does not | Staleness risk; semantic thresholds need tuning and monitoring |
| Quantization | Moderate, via memory headroom becoming batch size | Task-dependent quality loss that must be measured |
| Batch size to the SLO edge | Moderate | Per-request latency, directly and immediately |
| Spot or preemptible capacity | Moderate to large on the batch fleet only | Interruption handling; unusable for interactive traffic |
| Interactive and batch fleet split | Moderate, mostly by enabling other levers | Two fleets to operate and monitor |
| Max output token cap | Small usually, large on verbose workloads | Truncated answers if the cap is set below real need |
| Filling troughs with queued batch work | Large on owned capacity with spiky traffic | Needs genuinely deferrable work and a queue |
| Adding GPUs | Raises cost - it buys headroom, never efficiency | The usual response to a problem that was actually admission control |
Idle Time Is the Dominant Waste
On owned capacity, the biggest line item is almost never inefficient inference. It is GPUs that were on and doing nothing. Peak-sized capacity plus a daily traffic trough plus weekends produces an average utilisation far below what the hardware could sustain, and every idle hour is billed at the same rate as a saturated one.
The fix is not smaller capacity โ you still need peak. The fix is finding work to put in the troughs:
- Eval suites, regression runs, and judge grading.
- Embedding backfills and re-indexing after a chunking change.
- Bulk classification, enrichment, summarisation, and synthetic data generation.
- Anything with no user waiting on it.
Queue that work with a lower priority than interactive traffic, let the scheduler drain it whenever interactive demand drops, and preempt it instantly when demand returns. The batch fleet can run at a much larger batch size because it has no latency promise to keep. The pattern is ordinary Batch vs Stream processing applied to accelerators, and it is the cheapest utilisation improvement available because the capacity is already paid for.
Watch the corresponding metric. Requests per second tells you about demand; what you want on the dashboard is utilisation against provisioned capacity, plus queue depth and tokens in flight. See Performance Metrics.
Multi-Tenancy and Noisy Neighbours
Filling a fleet means sharing it, and sharing an accelerator is less isolated than sharing a web server. Contention shows up in ways that are easy to misdiagnose:
- One long-context tenant can starve everyone. KV cache is the shared scarce resource, so a handful of very long sequences consume the memory that would otherwise have held many short ones. Throughput for every other tenant drops even though nothing about their traffic changed.
- A burst of long generations stretches everyoneโs inter-token latency, because each decode step now covers more sequences. Tenants experience a latency regression they did not cause and cannot see.
- Fairness needs enforcing in tokens, not requests. A request-count quota treats a 30-token answer and a 3,000-token answer as equal, which they are not. Token-based limiting is covered at the system level in ChatGPT; the primitives are in Rate Limiting.
Practical defences: per-tenant caps on context length and max output tokens, per-tenant token budgets, priority classes so interactive traffic preempts batch, and dedicated replicas for any tenant whose traffic shape is genuinely hostile to the others. Attribute cost per tenant and per feature from day one โ you cannot have the conversation about who is expensive without the number.
Measure First, Then Optimise
Every figure in this page that matters is a measurement, not a constant: tokens per second per GPU at your SLO, concurrent sequences before KV memory is exhausted, your input-to-output token split, your real utilisation curve across a week, and cost attributed per feature. Published throughput numbers are measured on someone elseโs model, context length, batch size, and hardware, and they will not reproduce.
The failure mode is spending a sprint on quantization to win a moderate throughput improvement while the fleet sits at low average utilisation and half the traffic could have been served by a smaller model. Order of magnitude first, and always check where the money actually is before touching the thing that is most fun to tune.
When to Use
โ Do this capacity and cost work when:
- You are deciding between a hosted API and owned GPUs, and need a defensible number rather than a preference
- GPU or API spend is growing faster than usage, which almost always means context growth or a model creeping upward in a cheap path
- You are sizing a fleet before a launch and need a provisioning number with stated assumptions
- Latency is degrading under load and you need to know whether it is a capacity problem or an admission control problem
โ Do not spend time here when:
- You are on a hosted API at low volume - fix context length, caching, and model routing first, since those dominate
- You have not instrumented token usage per feature, in which case that is the task
- The feature is still a prototype and may not ship
- You are tempted to buy capacity to paper over unbounded concurrency, which just gets consumed the same way
Common Interview Questions
Q1: Hosted API or self-hosted GPUs โ how do you decide?
By computing both at my real volume and my real utilisation, because the two have structurally different cost shapes rather than different prices. Hosted is variable cost per token with no floor; owned GPUs are fixed cost per hour that bill identically whether saturated or idle. So I measure actual token volume split by input and output, price the hosted path at current published rates including retries and eval traffic, size a fleet from the memory-bound concurrency constraint, and divide fleet cost by tokens delivered at the utilisation I will realistically achieve rather than at capacity. Then I add the operational costs that never appear on the GPU invoice โ engineer time, on-call, warmup waste, over-provisioned headroom. Self-hosting wins at high steady well-utilised volume, and the common mistake is assuming a utilisation the traffic shape will never deliver.
Q2: Why is utilisation the thing you keep coming back to?
Because on fixed-cost capacity, cost per token is fleet cost divided by delivered tokens, and that relationship is linear. A fleet running at one tenth the utilisation of an identical fleet pays ten times as much per token for the same model on the same hardware. No kernel optimisation or quantization scheme recovers a factor of ten. It also explains why spiky traffic is the worst fit for owned capacity: you must provision for peak and you pay through the trough, so the quiet hours carry the peakโs cost. The corollary is that the cheapest optimisation available is usually queueing deferrable work โ evals, backfills, bulk classification โ into the troughs, because that capacity is already paid for.
Q3: What limits how many concurrent requests a GPU can serve, and why does that matter for cost?
KV cache memory. Weights take a fixed slice of device memory and activations take some working room; whatever remains has to hold one KV cache per in-flight sequence, growing with every token generated. So concurrency is a memory budget, not a compute budget, and when you size a fleet the memory-derived GPU count is usually larger than the throughput-derived one โ size on the larger or you under-provision. For cost, this means context length is a capacity lever and not only a token-price lever: doubling average context roughly halves how many sequences fit, which halves batch size and throughput per GPU, so the same traffic now needs about twice the fleet.
Q4: Walk me through sizing a fleet for 400 concurrent generations.
Two independent estimates, then take the larger. On the throughput side, the SLO converts users into a token rate โ if each stream needs a readable output rate, aggregate demand is concurrency times that rate โ and I divide by measured sustained throughput per GPU at the batch size that actually met the SLO, then inflate for the share of capacity prefill consumes. On the memory side, I divide concurrency by the measured number of sequences one GPU holds at my average context length before KV memory runs out. Memory normally binds, so that is the number, plus a replica for failure headroom and warm headroom sized to model load time. Every input there is a measurement on my own model and hardware; published throughput figures are measured on someone elseโs configuration and will not reproduce.
Q5: Your GPU bill doubled this quarter and traffic did not. Where do you look?
At context length first, then at routing, then at utilisation. Context growth is the usual culprit and it is invisible in request counts โ retrieval returning more chunks, a longer system prompt, more conversation history retained โ and on owned capacity it hits twice because it also shrinks batch size. Next, whether traffic that used to hit a cheap model is now hitting an expensive one, which happens quietly when a fallback or a default changes. Then whether provisioned capacity grew to cover a latency incident that was actually an admission control problem, since capacity added for that reason gets consumed the same way and never comes back. All three need per-feature token and cost attribution to diagnose, which is why that instrumentation goes in on day one rather than during the incident.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts