Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 16 min read

Tokens, Context Windows, and Cost Math - Complete Deep Dive

Prerequisites: How LLMs Actually Work, LLM APIs and SDKs Used in: Prompt Engineering, RAG End to End, Model Selection and Routing, Prompt and Semantic Caching Build it: Lesson 21 - Latency and Cost Engineering implements this as runnable, tested code you can execute offline.


What is a token?

A token is the unit a language model actually reads and writes. It is not a word and it is not a character β€” it is a subword fragment drawn from a fixed vocabulary of tens of thousands of pieces, learned from a training corpus. Common words are usually one token. Rare words, proper nouns, misspellings, and identifiers get chopped into several. The model never sees your text; it sees a list of integer ids.

This is the unit that every bill, every limit, and every latency budget in AI engineering is denominated in. If you are fuzzy about tokens, you will be fuzzy about cost, about why a request was rejected, and about why the model got a simple question wrong.

Real-world analogy: Freight shipping with a fixed set of crate sizes. You cannot ship β€œa chair” β€” you ship it in the crates available, and an awkwardly shaped item takes more crates than its size suggests. English prose packs efficiently because the crate sizes were chosen for it. A JSON blob, a Hindi sentence, or a base64 string packs badly and costs more crates for the same amount of meaning. And once packed, the shipper only sees crates. They cannot tell you how many letters are inside one.


Subword Tokenization and What It Costs You

Most production tokenizers come from the byte-pair encoding family. The construction is simple: start from bytes, repeatedly merge the most frequent adjacent pair into a new vocabulary entry, stop at a target vocabulary size. The result is a compression scheme tuned to the training distribution β€” frequent sequences become single tokens, everything else stays fragmented.

Four practical consequences, all of which show up in real systems:

Token density varies enormously by content type. The same amount of meaning costs different numbers of tokens depending on what it is written in.

Content Why it tokenizes the way it does Relative token cost per unit of meaning
Common English prose The vocabulary was largely fit to it Baseline
Technical English with rare terms Unusual words fragment into pieces Modestly above baseline
Source code Indentation, punctuation, and camel-case identifiers all split Noticeably above baseline
Non-Latin scripts Fewer merges were learned, so text falls back toward bytes Often several times baseline
Base64 blobs, UUIDs, hashes Effectively random, so no merges apply Worst case, near byte level

This has a direct product consequence that catches teams out: the same feature costs materially more per user in some languages than in English, and a long context window shrinks in effective capacity for those users. If you serve a multilingual product, measure token counts per language rather than assuming an English baseline holds.

The model cannot reliably count letters, and now you know why. Ask how many times β€œr” appears in β€œstrawberry” and the model is being asked about characters it never received. It got two or three opaque integer ids. Getting the answer right requires reasoning about a representation it cannot perceive, which is why letter counting, string reversal, and precise rhyming are unreliable in a way that feels absurd for a system that can write a working parser. The fix is not a better prompt β€” it is to do character-level work in code and let the model orchestrate.

Token counts are not portable across model families. Each family ships its own tokenizer, so identical text yields different counts. When you compare two providers’ prices per million tokens, you are comparing prices on two different units. Tokenize your actual traffic with each candidate’s tokenizer before concluding one is cheaper.

Boundaries are invisible but load-bearing. A trailing space, a different newline convention, or a changed indentation style alters the token sequence and therefore the output. This is one reason a prompt that β€œstopped working” after a harmless reformat really did change.


The Rule of Thumb, With the Caveat First

For English prose only, and only for rough planning:

Treat these as order-of-magnitude figures. They are wrong for code, wrong for non-English text, wrong for structured data, and they vary by tokenizer even within English. They are good enough to decide whether a design is plausible, and not good enough to enforce a limit.

# Estimate for a rough plan. Measure for anything that enforces a budget.
def estimate_tokens(text: str) -> int:
    return len(text) // 4

def count_tokens(text: str, tokenizer) -> int:
    return len(tokenizer.encode(text))   # the real number, per model family

If a budget check decides whether a request gets rejected, use the real tokenizer. A 25% estimation error is harmless in a spreadsheet and is a production outage in a context-limit guard.


The Context Window Is One Shared Budget

The context window is a hard ceiling on the total token count of a single call. The mistake is thinking of it as β€œhow much input I can send.” It is shared by everything in the request plus the space reserved for the answer.

flowchart LR
    W[Context window - one hard budget] --> S[System prompt - paid on every call]
    W --> H[Conversation history - grows every turn]
    W --> R[Retrieved context - you choose the amount]
    W --> O[Reserved output space - the max output setting]
    H --> P[Pressure point - trim or summarize]
    R --> P
    P --> F[Exceed the window and the call is rejected outright]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class W client
    class S,O data
    class H,R service
    class P async
    class F edge

Four allocations compete. The system prompt is fixed but paid every single call, so a bloated one is a permanent tax. History grows without bound unless you manage it. Retrieved context is the one you control most directly and usually the largest. And output space must be reserved β€” asking for a long answer when the input already fills the window fails, because the two come from the same pool.

Two more things are true of the window that the number on a spec sheet does not tell you. Filling it is not free: input tokens are billed and they add latency, so a bigger window is permission to spend more, not an invitation. And the usable window is smaller than the advertised one β€” recall of material buried in the middle of a very long context is widely reported to degrade relative to material at the start or end. Position your most important context deliberately rather than relying on the model to find it.

Bad β†’ Good β†’ Great: staying inside the window

Bad β€” stuff everything and truncate when it breaks. Concatenate the full history and every document you have, then chop the front off when the API complains. You silently delete the system prompt or the original task, and the model’s behaviour changes for reasons that never appear in a log.

Good β€” budget explicitly, then truncate oldest-first with pinning. Compute the budget before you send: window minus reserved output minus system prompt gives you what history and retrieval may share. Drop whole oldest turns rather than slicing mid-message, and pin anything that must survive β€” the system prompt and the original task statement. Predictable, and it fires as a decision you made rather than an error you caught.

Great β€” summarize the middle and retrieve instead of stuffing. Keep a rolling summary of old turns plus the last few verbatim, so a long conversation degrades into compressed memory rather than amnesia. And for documents, do not put the corpus in the prompt at all: index it and retrieve the handful of relevant chunks per question, which is the entire argument for RAG over long-context stuffing. Retrieval scales to a corpus no window will ever hold, costs less per request, and usually answers better because the model reads five relevant chunks instead of hunting through two hundred pages. Chunk sizing is its own problem β€” see Document Parsing and Chunking. The system-design treatment of conversation summarization at scale lives in Design ChatGPT.


Asymmetric Pricing: Input vs Output

Providers bill input and output tokens at different rates, and output is consistently the more expensive side β€” commonly by a multiple, not a few percent. The reason is mechanical: input tokens are processed in parallel in one prefill pass, while output tokens are generated one at a time, each requiring a full forward pass through the model.

Which side dominates your bill depends entirely on your workload shape, and this determines where optimisation effort belongs:

Workload Shape Cost driven by
RAG and document question answering Large retrieved input, short answer Input tokens, by a wide margin
Summarization Large input, moderate output Input tokens
Drafting, long-form generation, code writing Small brief, long output Output tokens
Agent loops with tool calls History regrows every step, many turns Input tokens, because history is resent each step
Classification and extraction Small in, tiny out Neither dominates much β€” volume is your problem, not shape

Agent loops deserve the warning. Each step resends the full history including every prior tool result, so a ten-step agent can pay for the same context ten times. That is the dominant cost of agentic systems and the reason step budgets and history compaction are not optional there.


Worked Cost Estimate

These unit prices are ASSUMED for the arithmetic below. They are not quotes from any provider. Real prices change frequently, differ per model, and must be looked up on the provider’s current pricing page before you commit to a number. What is durable here is the method and the relative sizes, not the dollar figures.

ASSUMED unit prices: input $0.50 per 1M tokens, output $1.50 per 1M tokens β€” a 3x output premium, which is a plausible order of magnitude.

Scenario A β€” a RAG support assistant

Component Tokens
System prompt 400
Conversation history retained 1,200
Retrieved chunks - 6 at 500 tokens each 3,000
User question 100
Input total 4,700
Output answer 300

Per user per month, assuming 40 assisted questions: 40 Γ— $0.0028 = $0.112, so about 11 cents per active user per month. At 10,000 monthly active users that is roughly $1,120 per month in model spend. Multiply by your P95 retry rate and any eval traffic before you quote it.

Scenario B β€” a drafting tool, same assumed prices

Input 600 tokens, output 1,500 tokens:

Nearly identical cost per request, mirror-image structure. In Scenario A, dropping from six retrieved chunks to three saves 1,500 input tokens, about 27% of the request. In Scenario B, capping output at 1,000 tokens instead of 1,500 saves about 29%. Now cross them over: trimming output in Scenario A saves around 5%, and trimming a couple of hundred input tokens in Scenario B saves under 4%. The same optimisation is a headline win in one workload and a rounding error in the other. Find your split before you optimise. The general estimation discipline is in Back-of-Envelope Estimation.


The Levers, Ranked

  1. Shorten the context. Almost always the largest and cheapest win, because it applies to every single request and requires no new infrastructure. Trim the system prompt, retrieve three excellent chunks instead of ten mediocre ones, compact history, and stop resending material the model does not need. Retrieval quality and cost improve together here, which is rare. Where output dominates, the equivalent first move is capping max_output_tokens and asking for terse answers.
  2. Use a smaller model for the easy traffic. Most production traffic is not hard. Route simple classification, routing, and extraction to a cheap small model and reserve the expensive one for cases that need it, with a measurable quality gate so you learn when the routing is wrong. The spread between tiers within a family is large enough that this often beats every other lever combined β€” see Model Selection and Routing.
  3. Cache. Repeated traffic is free traffic. Exact-match caching on a request key is trivial and catches more than you expect; provider-side prompt caching discounts a long stable prefix such as your system prompt; semantic caching catches near-duplicate questions at the cost of a similarity threshold you must tune. See Prompt and Semantic Caching and the general mechanics in Caching.
  4. Batch the work that can wait. Evals, backfills, bulk classification, and nightly enrichment have no user watching. Provider batch modes trade latency for a real discount, and moving that traffic off the interactive path also stops it competing for your rate limit.

Below those four, worth doing but smaller: set max_output_tokens on every call as a spend ceiling, keep n at 1, avoid re-asking the model for work you already have on disk, and instrument token usage per feature so you can see which surface is actually expensive. You cannot rank levers on a workload you have not measured β€” log prompt_tokens and completion_tokens on every call from day one, as covered in LLM Observability.


When to Use

βœ… Do this token and cost work when:

❌ Do not over-invest when:


Common Interview Questions

Q1: Why can a model write a working compiler but fail to count the letters in a word?

Because it never sees letters. Text is converted to subword token ids before the model reads anything, so a word like β€œstrawberry” arrives as a couple of opaque integers with no character structure exposed. Counting letters requires reasoning about a representation the model cannot perceive, whereas writing code is exactly the pattern-completion task it was trained on. The engineering lesson is to do character-level and arithmetic-level work in code and use the model to decide what to compute, not to compute it.

Q2: Your feature costs 3x more per user in one market than another. What is the likely cause?

Tokenization. Byte-pair vocabularies are fit largely to the training distribution, so text in scripts with fewer learned merges falls back toward byte-level encoding and consumes several times more tokens for the same meaning. The same sentence therefore costs more, and the context window holds proportionally less of the conversation. Check token counts per language with the real tokenizer, and consider a per-locale context budget rather than one global setting.

Q3: You have a 500-page manual and a large context window. Do you stuff it or retrieve?

Retrieve, in nearly every case. Stuffing pays input cost on the entire manual for every question, adds latency in proportion, and relies on the model locating one passage inside an enormous context β€” recall of buried middle content degrades noticeably. Retrieval sends a handful of relevant chunks, costs a fraction as much, is often more accurate because there is less to distract from, and keeps working when the manual grows past any window. Long context is the right answer when the whole document genuinely must be reasoned over at once, such as a contract consistency check, or when you have one throwaway question and no index.

Q4: Where does the money actually go in a RAG system versus a drafting assistant?

Opposite ends. RAG is input-dominated: a few thousand tokens of retrieved context against a few hundred tokens of answer, so input is typically the large majority of the bill even though output is priced higher per token. A drafting assistant inverts it β€” a short brief and a long generation, so output dominates despite being fewer total tokens in some cases. This matters because the correct optimisation differs: fewer and better retrieved chunks for the first, tighter output caps and terser instructions for the second. Measure your own split before choosing.

Q5: You need to cut model spend by half. What do you do, in order?

First cut context per request, since it applies to every call and needs no new infrastructure β€” trim the system prompt, retrieve fewer and better chunks, compact history, and cap output length. Second, route easy traffic to a smaller model behind a quality gate, because the price spread between tiers is usually the single largest lever available. Third, cache: exact-match on a request key, provider prompt caching on the stable prefix, semantic caching for near-duplicate questions. Fourth, move eval and backfill traffic to a batch mode so it is both discounted and off the interactive rate limit. All of this presumes per-feature token logging, because otherwise you are guessing which surface is expensive.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access