Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 20 min read

GPU Capacity and Cost Engineering - Complete Deep Dive

Stage 7 - Production Lesson 30 of 36

Prerequisites: Inference Serving, Tokens and Cost Math, Back-of-Envelope Estimation Used in: Latency Engineering, Model Selection and Routing, ChatGPT


What is GPU Cost Engineering?

GPU cost engineering is the practice of deciding how much accelerator capacity you need, whether you should own it at all, and which levers actually reduce the bill โ€” with measurements rather than intuition. It sits directly on top of Inference Serving: the serving stack determines how much work one GPU can do, and this page turns that number into a fleet size and a cost per million tokens.

Real-world analogy: Taxis versus a company car. A taxi charges per trip โ€” no trips, no bill, and the per-trip price never improves no matter how much you ride. A company car costs the same every month whether it is on the motorway or in the car park, so it only wins if you drive it constantly. Nobody buys a car for four trips a month, and nobody takes taxis for a 300-kilometre daily commute. Hosted API versus owned GPUs is exactly this decision, and utilisation is the only thing that decides it.


The Economics Are Structurally Different

flowchart TD
    subgraph Hosted["Hosted API - variable cost"]
        H1[Request arrives] --> H2[Pay per token consumed]
        H2 --> H3[No traffic means no spend]
        H3 --> H4[Cost tracks demand exactly and never amortises]
    end

    subgraph Owned["Self hosted GPUs - fixed cost"]
        O1[Provision capacity for peak] --> O2[Pay per GPU hour idle or saturated]
        O2 --> O3[No traffic still costs full price]
        O3 --> O4[Cost per token falls only as utilisation rises]
    end

    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef client fill:#f97316,stroke:#c2410c,color:#fff

    class H1 client
    class H2,H3,H4 service
    class O1 client
    class O2,O3,O4 data

A hosted API is variable cost per token. It is linear, it has no floor, and it has no volume-driven improvement unless you negotiate one. A self-hosted GPU is fixed cost per hour โ€” the meter runs identically at 4% utilisation and at 94%.

Three consequences that decide most architecture arguments:

  1. Utilisation is the entire game. Cost per token on owned capacity is fleet cost divided by delivered tokens. The numerator is fixed by your provisioning decision; the denominator is whatever traffic shows up. Double utilisation and you have halved cost per token without touching the model.
  2. Spiky traffic is the worst possible fit for owned capacity. You must provision for peak but you pay for the troughs, so a workload with a 10x peak-to-trough ratio is paying roughly 10x its marginal cost during the quiet hours. Business-hours-only internal tools are the classic mismatch: sized for 9am, idle for sixteen hours a day.
  3. Fixed cost punishes uncertainty. Hosted spend scales down with a failed launch. A reserved GPU fleet does not, and a one-year commitment on a product that might get cancelled in three months is a bet on the roadmap, not on the infrastructure.

This is also why the hybrid answer is so common and so sensible: own capacity for the steady base load, burst to a hosted API above it. The economics of each shape are matched to the traffic that suits them.


Memory Is the Binding Constraint

Before sizing anything, internalise what fills up on a serving GPU. Device memory holds three things:

Weights and overhead are constants. The KV cache is the variable, and it is the one that runs out. The consequence is the sentence to remember: concurrency is limited by KV cache memory, not by compute. You do not run out of arithmetic; you run out of room to remember what has already been said.

That is why context length is a capacity decision and not just a cost decision. Doubling average context roughly doubles KV cache per sequence, which roughly halves how many sequences fit, which halves your batch size and therefore your throughput per GPU. A feature that quietly grew its retrieved context from four chunks to twelve did not just raise the token bill โ€” it shrank the fleetโ€™s capacity, and that shows up as queueing and p99 rather than as a line item.


Worked Sizing Example

Every number in this section is an ASSUMPTION invented to make the arithmetic readable. None of it is a quote, a benchmark, or a spec for any real accelerator. Real serving throughput and real KV capacity must be measured on your own model, your own context lengths, and your own hardware, and hourly rates must be looked up on your providerโ€™s current pricing page. What transfers is the method and the shape of the result, never the figures.

The hourly rate appears below as R โ€” one GPU-hour, as a unit. Substitute your real rate at the end and the whole calculation scales linearly.

Step 1 - demand, from the SLO

Note that the SLO is what converts โ€œ400 usersโ€ into a token rate. Accept a slower stream and the same users need less capacity โ€” the dial from Inference Serving, applied to a budget.

Step 2 - supply, from a measurement

Step 3 - the memory check, which usually wins

Size on the larger of the two. Nine GPUs, because memory binds before compute does โ€” which is the normal outcome, and the reason a purely throughput-based estimate under-provisions. Add one replica for failure headroom and enough warm headroom to cover model load time, and you are provisioning 10 for a peak of 400.

Step 4 - cost per million tokens, where utilisation appears

ASSUMED average utilisation Delivered output tokens per month Cost per million output tokens
100% - unattainable 47,300 M 0.14 R
40% 18,900 M 0.35 R
10% 4,730 M 1.39 R

The ratio is the lesson. Cost per token is linear in utilisation, so a 10x utilisation gap is a 10x cost gap on identical hardware serving an identical model. No quantization scheme, batch size, or kernel optimisation will recover a 10x. Filling the fleet will.

The general discipline โ€” assume loudly, compute roughly, check the order of magnitude โ€” is in Back-of-Envelope Estimation.


Provisioning - Bad to Good to Great

Bad - provision for peak and leave it

Size the fleet for the worst hour of the week, run it flat out at that size forever, and treat the bill as the cost of doing business. This is the most common owned-capacity deployment and it is structurally wasteful: average utilisation ends up a fraction of peak, so most of what you pay for is idle silicon. It also fails in the other direction, because โ€œpeakโ€ was measured before the feature got popular.

Good - autoscale the fleet on demand

Scale replicas up and down with traffic. A real improvement, and it removes the overnight waste. Its limits are specific to accelerators: instances are scarce and slow to start, so reactive scaling arrives after the spike it was meant to absorb; warm headroom has to be paid for anyway to cover model load time; and scaling in aggressively means paying the load cost repeatedly when traffic oscillates. You also still have troughs โ€” smaller ones, but empty.

Great - a two-class fleet with queued backfill work

Provision the steady base load as owned capacity you keep busy, burst above it to a hosted API rather than to cold GPUs, and run a low-priority batch class that consumes whatever the interactive class is not using. Interactive traffic gets a modest batch size, a bounded queue, and preemption rights. Batch work โ€” evals, embedding backfills, bulk classification โ€” gets the largest batch memory allows, spot or preemptible capacity where available, and no latency promise. Utilisation rises toward what the hardware can actually sustain, and the peak is absorbed by the cost structure that handles peaks well.


The Crossover Analysis - A Method, Not a Verdict

The honest answer to โ€œis self-hosting cheaperโ€ is โ€œcompute it for your volume and your utilisation, and recompute it when either changes.โ€ What follows is the procedure, not a conclusion.

  1. Measure your actual token volume, split into input and output, per feature. If you cannot produce this number you are not ready to have the conversation. Instrument it as described in Tokens and Cost Math.
  2. Price the hosted path at current published rates for the model you would actually ship, using your real input-to-output split. Include retries, eval traffic, and internal testing โ€” teams routinely forget all three.
  3. Price the self-hosted path using the sizing above: fleet size from the memory-bound constraint, hourly rate from your provider, then divide by delivered tokens at your realistic utilisation, not at capacity.
  4. Add the costs that are not on the GPU invoice. Engineer time to operate the fleet, on-call burden, model loading and warmup waste, over-provisioned headroom, load testing, and the opportunity cost of the work those engineers are not doing. This is frequently larger than the hardware delta and is almost always left out of the spreadsheet that โ€œprovesโ€ self-hosting wins.
  5. Plot both against volume and find where the lines cross. Then ask the real question: how confident are you that you will sustain the volume and utilisation on the right-hand side of that crossing?

Two things are usually true. Self-hosting can be dramatically cheaper per token at high, steady, well-utilised volume with an open-weights model that is good enough for the task. And most teams do not have high, steady, well-utilised volume โ€” they have a launch curve, a weekday shape, and a model choice that is still moving. Break-even assumes the utilisation you will actually achieve, not the utilisation your capacity permits.

A useful middle path: prove the workload on a hosted API, measure the real traffic shape for a quarter, and only then move the steady base load onto owned capacity while bursting the peaks. You get the cost structure that fits each part of the curve. The model-side version of this decision is in Model Selection and Routing.


The Levers, Ranked

  1. Route easy traffic to a smaller model. The spread between model tiers is the largest single factor in most bills, and most production traffic is not hard. Classification, routing, extraction, and short factual answers rarely need your best model. Gate it with evals so you find out when the router is wrong. See Model Selection and Routing and Distillation and Small Models.
  2. Cut context length. Applies to every single request, needs no new infrastructure, and on owned capacity it pays twice โ€” less prefill compute and a smaller KV cache per sequence, which means a larger batch and more throughput per GPU. Retrieve three excellent chunks instead of twelve mediocre ones.
  3. Cache. Repeated traffic is free traffic. Exact-match caching is trivial, prefix and prompt caching discount a stable system prompt, and semantic caching catches near-duplicates at the price of a threshold you must tune. See Prompt and Semantic Caching and the mechanics in Caching.
  4. Quantize. On owned capacity this is a memory lever that converts into a throughput lever. Eval-gate it; a quality regression you did not measure is not a saving.
  5. Raise batch size to the edge of the SLO. Free throughput up to the point where p95 latency breaches the target, and not one step further.
  6. Use spot or preemptible capacity for batch work. Meaningful discounts for work that can be interrupted and retried. Never for the interactive path.
  7. Separate interactive from batch fleets. Mostly an enabler for the two levers above, and it stops a nightly bulk job from degrading chat latency.
  8. Cap maximum output tokens. A ceiling on the most expensive phase, and simultaneously a latency control.
Lever Typical magnitude Tradeoff
Route to a smaller model Largest available - often a multiple, not a percentage Router complexity; misroutes are quality regressions unless eval-gated
Cut context length Large, and it applies to every request Retrieval tuning effort; too aggressive and answers lose grounding
Caching Large where traffic repeats, negligible where it does not Staleness risk; semantic thresholds need tuning and monitoring
Quantization Moderate, via memory headroom becoming batch size Task-dependent quality loss that must be measured
Batch size to the SLO edge Moderate Per-request latency, directly and immediately
Spot or preemptible capacity Moderate to large on the batch fleet only Interruption handling; unusable for interactive traffic
Interactive and batch fleet split Moderate, mostly by enabling other levers Two fleets to operate and monitor
Max output token cap Small usually, large on verbose workloads Truncated answers if the cap is set below real need
Filling troughs with queued batch work Large on owned capacity with spiky traffic Needs genuinely deferrable work and a queue
Adding GPUs Raises cost - it buys headroom, never efficiency The usual response to a problem that was actually admission control

Idle Time Is the Dominant Waste

On owned capacity, the biggest line item is almost never inefficient inference. It is GPUs that were on and doing nothing. Peak-sized capacity plus a daily traffic trough plus weekends produces an average utilisation far below what the hardware could sustain, and every idle hour is billed at the same rate as a saturated one.

The fix is not smaller capacity โ€” you still need peak. The fix is finding work to put in the troughs:

Queue that work with a lower priority than interactive traffic, let the scheduler drain it whenever interactive demand drops, and preempt it instantly when demand returns. The batch fleet can run at a much larger batch size because it has no latency promise to keep. The pattern is ordinary Batch vs Stream processing applied to accelerators, and it is the cheapest utilisation improvement available because the capacity is already paid for.

Watch the corresponding metric. Requests per second tells you about demand; what you want on the dashboard is utilisation against provisioned capacity, plus queue depth and tokens in flight. See Performance Metrics.


Multi-Tenancy and Noisy Neighbours

Filling a fleet means sharing it, and sharing an accelerator is less isolated than sharing a web server. Contention shows up in ways that are easy to misdiagnose:

Practical defences: per-tenant caps on context length and max output tokens, per-tenant token budgets, priority classes so interactive traffic preempts batch, and dedicated replicas for any tenant whose traffic shape is genuinely hostile to the others. Attribute cost per tenant and per feature from day one โ€” you cannot have the conversation about who is expensive without the number.


Measure First, Then Optimise

Every figure in this page that matters is a measurement, not a constant: tokens per second per GPU at your SLO, concurrent sequences before KV memory is exhausted, your input-to-output token split, your real utilisation curve across a week, and cost attributed per feature. Published throughput numbers are measured on someone elseโ€™s model, context length, batch size, and hardware, and they will not reproduce.

The failure mode is spending a sprint on quantization to win a moderate throughput improvement while the fleet sits at low average utilisation and half the traffic could have been served by a smaller model. Order of magnitude first, and always check where the money actually is before touching the thing that is most fun to tune.


When to Use

โœ… Do this capacity and cost work when:

โŒ Do not spend time here when:


Common Interview Questions

Q1: Hosted API or self-hosted GPUs โ€” how do you decide?

By computing both at my real volume and my real utilisation, because the two have structurally different cost shapes rather than different prices. Hosted is variable cost per token with no floor; owned GPUs are fixed cost per hour that bill identically whether saturated or idle. So I measure actual token volume split by input and output, price the hosted path at current published rates including retries and eval traffic, size a fleet from the memory-bound concurrency constraint, and divide fleet cost by tokens delivered at the utilisation I will realistically achieve rather than at capacity. Then I add the operational costs that never appear on the GPU invoice โ€” engineer time, on-call, warmup waste, over-provisioned headroom. Self-hosting wins at high steady well-utilised volume, and the common mistake is assuming a utilisation the traffic shape will never deliver.

Q2: Why is utilisation the thing you keep coming back to?

Because on fixed-cost capacity, cost per token is fleet cost divided by delivered tokens, and that relationship is linear. A fleet running at one tenth the utilisation of an identical fleet pays ten times as much per token for the same model on the same hardware. No kernel optimisation or quantization scheme recovers a factor of ten. It also explains why spiky traffic is the worst fit for owned capacity: you must provision for peak and you pay through the trough, so the quiet hours carry the peakโ€™s cost. The corollary is that the cheapest optimisation available is usually queueing deferrable work โ€” evals, backfills, bulk classification โ€” into the troughs, because that capacity is already paid for.

Q3: What limits how many concurrent requests a GPU can serve, and why does that matter for cost?

KV cache memory. Weights take a fixed slice of device memory and activations take some working room; whatever remains has to hold one KV cache per in-flight sequence, growing with every token generated. So concurrency is a memory budget, not a compute budget, and when you size a fleet the memory-derived GPU count is usually larger than the throughput-derived one โ€” size on the larger or you under-provision. For cost, this means context length is a capacity lever and not only a token-price lever: doubling average context roughly halves how many sequences fit, which halves batch size and throughput per GPU, so the same traffic now needs about twice the fleet.

Q4: Walk me through sizing a fleet for 400 concurrent generations.

Two independent estimates, then take the larger. On the throughput side, the SLO converts users into a token rate โ€” if each stream needs a readable output rate, aggregate demand is concurrency times that rate โ€” and I divide by measured sustained throughput per GPU at the batch size that actually met the SLO, then inflate for the share of capacity prefill consumes. On the memory side, I divide concurrency by the measured number of sequences one GPU holds at my average context length before KV memory runs out. Memory normally binds, so that is the number, plus a replica for failure headroom and warm headroom sized to model load time. Every input there is a measurement on my own model and hardware; published throughput figures are measured on someone elseโ€™s configuration and will not reproduce.

Q5: Your GPU bill doubled this quarter and traffic did not. Where do you look?

At context length first, then at routing, then at utilisation. Context growth is the usual culprit and it is invisible in request counts โ€” retrieval returning more chunks, a longer system prompt, more conversation history retained โ€” and on owned capacity it hits twice because it also shrinks batch size. Next, whether traffic that used to hit a cheap model is now hitting an expensive one, which happens quietly when a fallback or a default changes. Then whether provisioned capacity grew to cover a latency incident that was actually an admission control problem, since capacity added for that reason gets consumed the same way and never comes back. All three need per-feature token and cost attribution to diagnose, which is why that instrumentation goes in on day one rather than during the incident.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access