How LLMs Actually Work - Complete Deep Dive
Prerequisites: The AI Engineer Role Used in: Tokens and Cost Math, Prompt Engineering, Hallucination and Grounding, Fine-Tuning
What is an LLM?
An LLM is a function. You give it a sequence of tokens and it returns a probability distribution over which token comes next - a score for every single token in its vocabulary. That is the whole interface. Everything else, including chat, reasoning, and tool calling, is built on top of that one operation being repeated.
Write it as a signature and the mystery drains out:
# The entire model, conceptually. Deterministic, stateless, one step.
def model(tokens: list[int]) -> dict[int, float]:
"""Returns P(next_token | tokens) over the full vocabulary. Sums to 1."""
Real-world analogy: think of a musician who has heard an enormous amount of music and can play the single most plausible next note for any passage you hum. Do that repeatedly and you get a melody that sounds right. The musician is not consulting a score, and is not checking whether the melody is factually about anything. They are producing what fits. An LLM does this with text, and the βfitsβ judgement is extremely good, which is exactly why it is easy to mistake for knowing.
This page exists so that model behaviour stops surprising you. Every section ends with the engineering consequence, because in production the consequence is the part you get paged for.
The Core Mechanic - Next-Token Prediction
The model is stateless and produces one token per forward pass. To generate a sentence, the system calls it in a loop, appending each chosen token to the input and running it again. This is called autoregressive generation.
flowchart LR
A[Prompt text] --> B[Tokenizer<br>text becomes token ids]
B --> C[Forward pass<br>one pass equals one token]
C --> D[Logits<br>one raw score per vocabulary token]
D --> E[Sampler<br>temperature and top-p applied]
E --> F[Pick one token]
F --> G[Append to the sequence]
G --> C
F --> H[Detokenize and stream to the client]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A,H client
class B,E edge
class C,F,G service
class D data
Three facts follow directly from that loop, and all three are load-bearing:
- Output length drives latency, not input length. The input is processed in one largely parallel pass - the prefill. Output tokens are generated strictly one at a time - the decode. A 100-token answer costs roughly 100 sequential forward passes. This is why βbe conciseβ is a latency optimization and not just a style note.
- Streaming is free and non-streaming is a deliberate delay. Tokens exist one at a time anyway, so streaming them is the natural behaviour. Buffering the full response means your users wait for the slowest token. Two separate metrics matter: time to first token, and tokens per second after that. See Latency Engineering.
- The model has no memory between calls. Conversation history is re-sent on every request. The illusion of memory is your application replaying the transcript, which is why long conversations get progressively slower and more expensive. See Tokens and Cost Math and Agent Memory and State.
The system-design implications of serving this loop at scale - batching, KV cache, GPU utilization - are covered in Inference Serving and at the architecture level in ChatGPT.
The Three Training Stages
A chat model is not produced in one step. It is produced by three stages with different data, different objectives, and different things they buy you.
flowchart LR
A[Raw text at web scale] --> B[Pretraining<br>predict the next token]
B --> C[Base model<br>fluent and knowledgeable<br>does not follow instructions]
C --> D[Supervised fine tuning<br>curated prompt and response pairs]
D --> E[Instruct model<br>follows instructions in a chat format]
E --> F[Preference tuning<br>rank competing answers]
F --> G[Aligned chat model<br>tone - refusals - helpfulness]
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A data
class B,D,F service
class C,E async
class G service
| Stage | Data | Objective | What it buys you | What it cannot fix |
|---|---|---|---|---|
| Pretraining | Enormous unlabelled text corpora | Predict the next token | Grammar, world knowledge, reasoning patterns, code, translation - nearly all raw capability | Behaviour. The result continues text rather than answering you |
| Supervised fine-tuning | Curated prompt and response pairs written or vetted by humans | Imitate the demonstrated responses | Instruction following, the chat role format, output style, task shape | Knowledge it never learned in pretraining. Preference between two acceptable answers |
| Preference tuning | Pairs of responses with a human or model judgement of which is better | Increase the score of preferred responses | Helpfulness, tone, refusal behaviour, less rambling, fewer unsafe completions | Factual accuracy. Ranking preference is not verifying truth |
The engineering consequence is a diagnosis tool. When a model misbehaves, ask which stage owns the problem:
- Wrong facts - a pretraining knowledge gap. Fix by supplying context at inference time, which is retrieval, not training. See RAG End to End.
- Wrong format or ignored instructions - an SFT-shaped problem. Fix with prompting, schema constraints, or light fine-tuning. See Structured Outputs.
- Wrong tone, over-refusal, or unhelpful hedging - preference-tuning behaviour. Fix with system prompts or a differently tuned model. See Preference Tuning.
Getting this mapping right saves enormous amounts of wasted effort. Teams routinely try to fine-tune away a knowledge gap, which is the most expensive possible way to fail at it.
Base Models vs Instruction-Tuned Models
| Β | Base model | Instruction-tuned or chat model |
|---|---|---|
| Trained through | Pretraining only | Pretraining plus SFT plus preference tuning |
| Given βWrite a haiku about cachingβ | May continue with more prompt-like text, or a list of similar exercises | Writes the haiku |
| Interface | Raw text continuation | Structured messages with system, user, and assistant roles |
| Good for | Fine-tuning starting point, research, raw completion tasks | Essentially all product work |
| Failure mode | Feels broken until you write few-shot prompts | Can be overly chatty or hedge when you wanted terse output |
If you download an open-weights model and it behaves strangely, check whether you picked the base variant rather than the instruct variant. This is a routine first-week mistake and costs people a day.
Attention - The Model Decides Which Earlier Tokens Matter
For each token it processes, the model computes how relevant every earlier token is to that position, then builds its representation mostly out of the relevant ones. That mechanism is attention, and it is what replaced the older approach of squeezing all previous context into one fixed summary vector.
Intuitive analogy: you are reading a long contract and hit the word βitβ in clause 40. To resolve βitβ you do not reread all 39 clauses equally. You skim back, land on the two clauses that define the subject, and ignore the rest. Attention is that skim, done numerically, for every token, in parallel. When resolving a pronoun the model weights the noun it refers to heavily; when closing a bracket it weights the opening bracket.
Engineering consequences:
- Long context is not free. Every token can attend to every earlier token, so attention work grows superlinearly with sequence length. Doubling your context more than doubles the compute. Modern implementations reduce the constant factor substantially, but the shape remains: stuffing the window is a real cost decision, not a free upgrade.
- Position inside the prompt matters. Attention is learned, not uniform, so content at the start and end of a long context tends to be used more reliably than content buried in the middle. Put instructions and the highest-relevance retrieved chunks at the edges. Ranking quality therefore matters even when everything fits - see Hybrid Search and Reranking.
- More context can lower accuracy. Ten irrelevant chunks give attention ten opportunities to be distracted. Retrieving fewer, better chunks usually beats retrieving more.
Sampling - What Temperature and Top-p Actually Do
The forward pass produces logits, which become a probability distribution. Sampling is the separate step that picks one token from that distribution, and it is where randomness enters. The model itself did not choose.
Suppose the distribution for the next token looks like this:
" fast" 0.50
" quick" 0.25
" cheap" 0.15
" purple" 0.02
...thousands more tokens with tiny probabilities
| Control | What it does | Effect |
|---|---|---|
| Temperature | Divides the logits before normalizing. Low values sharpen the distribution, high values flatten it | Near 0 the top token dominates and output is repetitive but stable. High values give the long tail real probability mass, which reads as creative and eventually as incoherent |
| Top-p (nucleus) | Keeps only the smallest set of tokens whose probabilities sum to p, renormalizes, then samples | Truncates the tail adaptively. At p equal to 0.9 the example above samples from the first three tokens and β purpleβ is impossible |
| Top-k | Keeps only the k highest-probability tokens | Same idea with a fixed cut instead of a probability mass cut. Less adaptive than top-p |
Practical guidance: for extraction, classification, structured output, and tool calling, use the lowest temperature available and rely on schema constraints. For drafting and brainstorming, moderate temperature plus top-p around 0.9 to 0.95 is a reasonable default. Tuning both at once mostly makes behaviour hard to reason about - move one.
Why the same prompt gives different answers
Four independent reasons, and only the first is the one people think of:
- Sampling is stochastic. Above zero temperature you are drawing from a distribution, so repeats differ.
- Temperature 0 is not a determinism guarantee. It makes the choice greedy, but floating-point addition is not associative and GPU kernels sum in whatever order the batch shape dictates. Two nearly tied tokens can flip between runs.
- Batching changes the arithmetic. Your request is batched with other tenantsβ requests, so the kernel shapes differ between calls. Same input, slightly different numerics.
- Hosted models change under you. Providers update serving stacks and model versions. Pin versions where the API allows it, and treat provider changes as a dependency upgrade that needs a re-run of your evals.
The consequence is architectural: never build a correctness assumption on byte-identical output. Validate structure with a schema, assert on semantics rather than exact strings, and make retries idempotent - see Idempotency. Snapshot tests over raw model output will fail for reasons unrelated to quality.
Why LLMs Hallucinate
Hallucination is not a bug on top of the model. It is the training objective working exactly as specified.
The objective rewards producing the most plausible continuation. It never contains a term for βis this trueβ and never contains a term for βsay you do not know.β A confident wrong citation and a confident correct citation look identical to the loss function - both are fluent, well-formed text of the right shape. So the model optimizes for text that looks like a correct answer, and most of the time that coincides with being correct, because in the training data plausible text usually was correct. When it does not coincide, you get a hallucination, delivered in the same confident register as everything else.
Preference tuning makes this partly worse before it makes it better: a confident answer often gets ranked above a hedge, which rewards sounding sure.
Bad: instruct the model not to hallucinate. βOnly state facts you are certain aboutβ moves the number slightly and leaves you with no defence, because the model has no reliable introspective access to its own certainty.
Good: ground it. Retrieve real source material, put it in the context, and instruct the model to answer only from that material and to say when the material is insufficient. This works because you have changed the task from recall to reading comprehension, which is far more reliable.
Great: ground it, then verify it. Require span-level citations back into the provided context, programmatically check that each cited span exists in what you supplied, and drop or flag claims that fail the check. Add an eval set that specifically includes unanswerable questions so you measure the abstention rate, not just the accuracy rate. Log-probabilities give you a weak confidence signal for routing borderline cases to review. Full treatment in Hallucination, Grounding, and Citations.
Why It Cannot Count Characters or Do Arithmetic
Two separate causes, and both are structural.
Tokenization hides characters. The model never sees letters. It sees token ids, where a token might be a whole word, a fragment, or a single character depending on how common the string is. Asking how many times a letter appears in a word is asking the model to report on information that was discarded before the first layer ran. It answers from the statistics of similar questions in training, which is guessing. Detail in Tokens, Context Windows, and Cost Math.
Arithmetic has no working memory. Multi-digit multiplication is an algorithm with intermediate state. The model produces one token per pass with no scratch space other than the tokens it has already emitted. Small numbers are memorized patterns from training. Larger ones require actually running the algorithm, and the model approximates instead - producing answers with the right digit count and right magnitude that are simply wrong.
The engineering consequence is the same in both cases and it is not a prompt: give it a tool. Route arithmetic to a calculator, counting to code, date math to a date library, and lookups to a query. Writing the expression is a language task and the model is good at it. Evaluating the expression is a computation task, so let a computer do it. See Tool Calling. Asking the model to show its working helps a little, because emitted tokens act as external scratch space, but a tool call is correct rather than merely better.
Engineering Consequences Cheat Sheet
| Fact about the model | What it means for your system |
|---|---|
| One forward pass per output token | Latency scales with output length. Cap max tokens and ask for brevity |
| Stateless between calls | You re-send history every turn. Summarize or window long conversations |
| Distribution plus sampler | Output is not reproducible. Validate schemas, never assert exact strings |
| Attention cost grows with context | A full context window is a cost and quality decision, not a free upgrade |
| Position within context matters | Put instructions and best evidence at the edges. Rerank before you truncate |
| Objective rewards plausibility | Grounding and verification are mandatory, not optional polish |
| Characters and digits are invisible | Delegate counting and arithmetic to tools |
| Knowledge is frozen at pretraining | Anything current must arrive through retrieval |
When to Use
β Reach for an LLM when:
- The task is language-shaped - summarizing, extracting, classifying, rewriting, translating, generating text or code
- The input is messy and unstructured in ways rules cannot enumerate
- An approximately correct answer is useful and a wrong one is recoverable or reviewable
- You can define what correct looks like well enough to build an eval set
β Do not reach for an LLM when:
- The task is exact computation - arithmetic, counting, sorting, aggregation. Use code
- A deterministic rule or a database query already answers it. A regex does not hallucinate
- Being silently wrong is unacceptable and no verification layer is possible
- The latency or cost budget cannot absorb a model call, and a cheaper classifier would do
Common Interview Questions
Q1: Explain what an LLM does in one sentence, then in three.
One sentence: it maps a sequence of tokens to a probability distribution over the next token. Three: that function is called repeatedly, with each chosen token appended to the input, which is autoregressive generation. Randomness lives in the sampler that picks from the distribution, not in the model, which is deterministic given identical inputs and identical numerics. Everything product-facing - chat, reasoning traces, tool calls - is that loop plus formatting conventions the model was fine-tuned to follow.
Q2: Why do LLMs hallucinate, and can you train it away?
They hallucinate because the objective rewards plausible continuations and contains no term for truth or for admitting ignorance. A fluent wrong citation and a fluent right one are indistinguishable to the loss. So it cannot be trained away in general - you can shift the rate but not eliminate the failure mode. The engineering answer is to change the task: supply the source material and turn recall into reading comprehension, require citations into that material, verify the cited spans programmatically, and measure abstention on deliberately unanswerable questions.
Q3: What is the difference between temperature and top-p?
Both shape the distribution before sampling, but differently. Temperature rescales the logits - low sharpens toward the top token, high flattens and gives the tail real mass. Top-p truncates: it keeps the smallest set of tokens whose probability sums to p and discards the rest, so the number of candidates adapts to how confident the model is at that step. I usually move one, not both: near-zero temperature for extraction and tool calling, and moderate temperature with top-p around 0.9 for generative work.
Q4: Why can a model write a working sort function but not reliably count letters in a word?
Because they are different kinds of task relative to the architecture. Writing the function is a language task - the pattern is well represented in training and the output is text. Counting letters requires character-level access that tokenization destroyed before the first layer, and multi-digit arithmetic requires intermediate working state the architecture does not have, since each pass emits one token with no scratch space. So the correct engineering move is not a better prompt, it is a tool call - let the model write the expression and let code evaluate it.
Q5: Same prompt, same temperature 0, different outputs. What is going on?
Temperature 0 makes the selection greedy, but it does not make the arithmetic bit-identical. Floating-point addition is not associative, and GPU kernels sum in an order that depends on batch shape, so a request batched alongside different traffic produces slightly different logits. When two candidate tokens are nearly tied, that is enough to flip the choice, and one flipped token changes everything after it. Hosted providers also update serving stacks and model versions underneath you. So I design for it: pin versions where possible, validate structure with a schema rather than string equality, and re-run the eval suite on any provider change.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts